🧠
Choose a Model
Pick a language model to preload and chat with. Every option runs entirely in your browser — nothing is uploaded anywhere.
Qwen3-0.6B
WebGPU ~340 MB q4f16
More capable, general-purpose chat. Requires a WebGPU-capable browser — Chrome or Edge on desktop.
Gemma 3 1B
WebGPU ~600 MB q4f16
Google's compact, newest-generation model — balanced quality and speed. Requires a WebGPU-capable browser — Chrome or Edge on desktop.
Phi-4-mini
WebGPU ~2.2 GB q4f16
Microsoft's largest option here — the biggest download by far, but the most capable answers. Requires a WebGPU-capable browser — Chrome or Edge on desktop.
SmolLM2-135M-Instruct
WASM CPU ~134 MB int8
Much smaller and faster to load, runs on any modern browser via WebAssembly — no GPU needed, simpler answers. The same model used by the AI Inference Benchmark.
Starting…
Qwen3-0.6B · q4f16 · WebGPU · Powered by WebLLM · Model by Qwen

About this tool

Supported models