Initializing, please wait a moment

Example conversation - written ahead of time so you can see the format. Your own chats are generated on your device after you load a model.
You
Turn these notes into a short status update: report draft done, waiting on two reviews, demo moved to Friday.
Example reply
Status update: the report draft is finished and is now with two reviewers. The demo has moved to Friday. Next step: collect the review feedback and finalize the report.

Saves the current conversation as a text file on your device.

Everything runs on your device. Your messages are not uploaded and there is no account. The model downloads once (0.3 GB and up, size shown in the picker) the first time and is cached by your browser for next time. A WebGPU browser is required for the model picker; Chrome with built-in AI can chat without it.

Private AI Chat


Chat with an AI model that runs entirely on your device, in your browser. It suits quick drafting, brainstorming, or asking a question you would rather keep off a server.

The AI model runs in your browser on your device - messages are not uploaded
Pick a model and chat on your device; nothing is uploaded.

How it works

How Private AI Chat works: pick a model and click Load model - the first time the weights download to your browser (0.3 GB and up, stated per model) and cache for next time; after that chat runs entirely on your device with no message upload and no account (WebGPU required). In Chrome without WebGPU, the page falls back to Chrome's built-in AI when that model is already on the device - still with no upload and no extra download.


Choosing a model

Private AI Chat - four of the available models: TinyLlama 1.1B, Qwen2.5 0.5B, Llama 3.2 1B, Llama 3.2 3B
Pick a smaller model for speed or a larger one for stronger answers.

Choosing a model in Private AI Chat: Qwen2.5 0.5B is the smallest first download (about 0.3 GB) and the fastest way to start; larger ones like Llama 3.2 3B or Phi-3.5 mini give stronger answers but take a bigger first download - switch models at any time. The sizes below were measured from the model files themselves.

ModelApprox. first download
Qwen2.5 0.5Babout 0.3 GB
TinyLlama 1.1Babout 0.6 GB
Llama 3.2 1Babout 0.7 GB
Qwen2.5 1.5Babout 0.8 GB
Llama 3.2 3Babout 1.7 GB
Phi-3.5 mini 3.8Babout 2 GB

Key features

Key features of Private AI Chat: streaming replies generated on your hardware and in-session history that stays in this tab only - nothing is saved or synced, and messages are never uploaded.

  • Streaming replies: answers stream in token by token, generated on your own hardware.
  • In-session history: the chat history stays in this tab for the session and is not saved or synced anywhere.
  • Answer presets: pick General assistant, Writing helper, Explain step by step, or Brainstorm ideas - changing the preset restarts the conversation so the new instruction applies from the first turn.
  • Transcript download: save the current conversation as a .txt file, built on your device - nothing is uploaded or synced.
  • Built-in AI fallback: in Chrome without WebGPU, the chat can run on Chrome's built-in AI when that model is already installed - zero extra download; the page never triggers Chrome's own model download.

Working with a document instead of an open-ended chat? Chat with PDF answers questions grounded in a PDF you open, with page-number citations - it reads the document on your device the same private way.

Drafting fiction rather than chatting? The AI story generator turns a premise, genre and length into flash fiction or a story opening on the same on-device models - it shares the model cache with this page, so a model downloaded here loads instantly there.

← Back to utility tools

Related tools:

Tags: #utility

Related guides:

Loading reviews...

Frequently Asked Questions

Does my chat leave my device?

No. The model runs in your browser and your messages stay on your device.

Why is there a large download the first time?

The AI model weights (from about 0.3 GB, shown per model in the picker) download once from the model registry and are then cached by your browser, so later visits start faster. The model is not bundled with the page.

Which browsers work?

The model picker needs WebGPU - the latest Chrome or Edge on desktop, or Chrome on Android. In Chrome without WebGPU, the chat can still run on Chrome's built-in AI when that model is already installed. Other browsers without either show a notice instead; there is no CPU fallback for the downloadable models.

Which model should I pick?

Qwen2.5 0.5B is the smallest download (about 0.3 GB) and the fastest way to start; Llama 3.2 3B and Phi-3.5 mini answer better but download more. You can switch models at any time.

Can I save the conversation?

Yes - the Download transcript button saves the current conversation as a .txt file, built on your device. Nothing is uploaded or synced; reloading the page still clears the in-tab history.