Llama 3.2 1BQ4_K_M
1B params · 800 MB
FeaturedIdeal for autocomplete, classification, and low-latency chat. Runs in-browser with WebGPU.
Available quants
Q4_K_MQ5_K_MQ8_0
fastbrowser-readytinyedge
Llama 3.2 3BQ4_K_M
3B params · 1.9 GB
Versatile 3B model for on-device chat and reasoning. Fits comfortably in mid-range mobile VRAM.
Available quants
Q4_K_MQ5_K_MQ8_0
fastbrowser-readyinstruction
Llama 3.1 8BQ4_K_M
8B params · 4.9 GB
Balanced 8B with a 128k context window. Great for document summarisation and tool-use.
Available quants
Q4_K_MQ5_K_MQ6_KQ8_0
long-contextinstructiondesktop
Llama 3.3 70BQ4_K_M
70B params · 43.0 GB
FeaturedFrontier-class open model. Requires multi-GPU desktop or mesh dispatch. Not browser-eligible.
Available quants
Q4_K_MQ5_K_M
frontiermesh-onlyhigh-qualitylong-context
Phi-3.5 MiniQ4_K_M
3.8B params · 2.3 GB
FeaturedOutstanding reasoning at 3.8B. 128k context window. Runs on mid-range consumer GPUs.
Available quants
Q4_K_MQ5_K_MQ8_0fp16
long-contextreasoningbrowser-readyinstruction
Phi-4 14BQ4_K_M
14B params · 8.5 GB
FeaturedMicrosoft's latest Phi-4 excels at STEM, coding and multi-step reasoning. Strong math performance.
Available quants
Q4_K_MQ5_K_MQ8_0
reasoningcode generationmathdesktop
Qwen 2.5 0.5BQ8_0
0.5B params · 500 MB
Ultra-lightweight multilingual model for edge deployment. Fits on any device with a WebGPU runtime.
Available quants
Q4_K_MQ8_0
tinyfastmultilingualbrowser-readyedge
Qwen 2.5 7BQ4_K_M
7B params · 4.7 GB
Strong multilingual and coding capabilities. Best run on desktop or mesh nodes.
Available quants
Q4_K_MQ5_K_MQ8_0
multilingualcode generationdesktop
Qwen 2.5 Coder 7BQ4_K_M
7B params · 4.7 GB
FeaturedSpecialised coding variant of Qwen 2.5. Supports fill-in-middle (FIM) for IDE-style completions.
Available quants
Q4_K_MQ5_K_MQ8_0
code generationfill-in-middledesktop
Qwen 2.5 72BQ4_K_M
72B params · 44.0 GB
FeaturedTop-tier multilingual model with 128k context. Rivals GPT-4o on many benchmarks. Mesh-dispatched.
Available quants
Q4_K_MQ5_K_M
frontiermultilinguallong-contextmesh-only
Mistral 7B v0.3Q4_K_M
7B params · 4.1 GB
Reliable general-purpose model with strong instruction following. Wide hardware support.
Available quants
Q4_K_MQ5_K_MQ8_0
generaldesktopinstruction
Mistral Nemo 12BQ4_K_M
12B params · 7.2 GB
Mistral × NVIDIA 12B with 128k context. Excellent multilingual performance and function-calling.
Available quants
Q4_K_MQ5_K_MQ8_0
long-contextmultilingualinstructiondesktop
Mixtral 8×7B MoEQ4_K_M
47B (MoE) params · 26.0 GB
Mixture-of-Experts with 8×7B experts, activating 13B per token. High quality at competitive speed.
Available quants
Q4_K_MQ5_K_M
moemesh-onlyhigh-qualitymultilingual
Gemma 2 2BQ4_K_M
2B params · 1.6 GB
Google's compact 2B model trained with knowledge distillation. Punches above its weight.
Available quants
Q4_K_MQ8_0
fastbrowser-readyinstructionedge
Gemma 2 9BQ4_K_M
9B params · 5.5 GB
Highly capable 9B from Google DeepMind. Competitive with models twice its size on many benchmarks.
Available quants
Q4_K_MQ5_K_MQ8_0
high-qualitydesktopinstruction
Gemma 2 27BQ4_K_M
27B params · 16.5 GB
Google's most capable open Gemma model. Excellent reasoning and instruction following.
Available quants
Q4_K_MQ5_K_M
high-qualitymesh-onlyreasoning
DeepSeek R1 7BQ4_K_M
7B params · 4.7 GB
FeaturedChain-of-thought reasoning specialist distilled from DeepSeek-R1. Excels at math and coding.
Available quants
Q4_K_MQ5_K_MQ8_0
reasoninglong-contextmathcode generationdesktop
DeepSeek R1 14BQ4_K_M
14B params · 8.7 GB
Larger R1 distillation with stronger multi-step reasoning. Best for complex analysis tasks.
Available quants
Q4_K_MQ5_K_M
reasoninglong-contextmathcode generationmesh-only
DeepSeek Coder V2 16BQ4_K_M
16B (MoE) params · 9.8 GB
State-of-the-art coding MoE model. Supports 338 programming languages and 160k context.
Available quants
Q4_K_MQ5_K_M
code generationfill-in-middlelong-contextmesh-only
Command R 35BQ4_K_M
35B params · 20.5 GB
Optimised for retrieval-augmented generation (RAG). Grounding, citations, and tool-use built in.
Available quants
Q4_K_MQ5_K_M
raglong-contextmultilingualmesh-only
StarCoder2 3BQ8_0
3B params · 3.2 GB
Lightweight code generation with fill-in-middle. Trained on 600+ languages. Browser-eligible.
License
BigCode OpenRAIL-M
Available quants
Q4_K_MQ8_0
code generationfill-in-middlebrowser-ready
StarCoder2 15BQ4_K_M
15B params · 9.1 GB
High-quality 15B code model with FIM. Best performance on competitive programming tasks.
License
BigCode OpenRAIL-M
Available quants
Q4_K_MQ5_K_MQ8_0
code generationfill-in-middledesktop
Nomic Embed v1.5Q8_0
137M params · 140 MB
High-quality text embeddings at near-zero cost. Runs in any browser, sub-5ms per batch.
embeddingsfastbrowser-readyedge
BGE-M3Q8_0
570M params · 580 MB
FeaturedState-of-the-art multilingual embedding model (100+ languages). Dense + sparse + ColBERT in one.
Available quants
Q4_K_MQ8_0fp16
embeddingsmultilingualbrowser-readyrag
E5-Mistral 7BQ4_K_M
7B params · 4.3 GB
Instruction-tuned LLM repurposed for high-quality long-context embeddings. Best for RAG pipelines.
Available quants
Q4_K_MQ8_0
embeddingslong-contextragdesktop
LLaVA 1.6 7BQ4_K_M
7B params · 4.9 GB
Visual instruction-tuned model. Accepts images + text. Suitable for document and chart analysis.
Available quants
Q4_K_MQ5_K_M
visionmultimodaldesktop
Moondream 2Q8_0
1.86B params · 1.9 GB
Tiny vision-language model designed for edge deployment. Runs in-browser for image Q&A.
Available quants
Q4_K_MQ8_0
visionmultimodalbrowser-readytiny