Developer Tool · Models

Model Library

Browse pre-quantized models, check hardware requirements, and deploy directly to local or mesh nodes with one click.

Showing 27 of 27 models
Llama 3.2 1BQ4_K_M
1B params · 800 MB
Featured

Ideal for autocomplete, classification, and low-latency chat. Runs in-browser with WebGPU.

VRAM
1 GB
RAM
2 GB
Browser tok/s
~45
Desktop tok/s
~90
Mobile tok/s
~18
Context
8k tokens
License
Meta Llama 3
Available quants
Q4_K_MQ5_K_MQ8_0
fastbrowser-readytinyedge
Llama 3.2 3BQ4_K_M
3B params · 1.9 GB

Versatile 3B model for on-device chat and reasoning. Fits comfortably in mid-range mobile VRAM.

VRAM
3 GB
RAM
4 GB
Browser tok/s
~28
Desktop tok/s
~72
Mobile tok/s
~10
Context
8k tokens
License
Meta Llama 3
Available quants
Q4_K_MQ5_K_MQ8_0
fastbrowser-readyinstruction
Llama 3.1 8BQ4_K_M
8B params · 4.9 GB

Balanced 8B with a 128k context window. Great for document summarisation and tool-use.

VRAM
6 GB
RAM
8 GB
Browser tok/s
~11
Desktop tok/s
~50
Context
131k tokens
License
Meta Llama 3
Available quants
Q4_K_MQ5_K_MQ6_KQ8_0
long-contextinstructiondesktop
Llama 3.3 70BQ4_K_M
70B params · 43.0 GB
Featured

Frontier-class open model. Requires multi-GPU desktop or mesh dispatch. Not browser-eligible.

VRAM
47 GB
RAM
63 GB
Desktop tok/s
~14
Context
128k tokens
License
Meta Llama 3
Available quants
Q4_K_MQ5_K_M
frontiermesh-onlyhigh-qualitylong-context
Phi-3.5 MiniQ4_K_M
3.8B params · 2.3 GB
Featured

Outstanding reasoning at 3.8B. 128k context window. Runs on mid-range consumer GPUs.

VRAM
3 GB
RAM
4 GB
Browser tok/s
~22
Desktop tok/s
~68
Mobile tok/s
~8
Context
128k tokens
License
MIT
Available quants
Q4_K_MQ5_K_MQ8_0fp16
long-contextreasoningbrowser-readyinstruction
Phi-4 14BQ4_K_M
14B params · 8.5 GB
Featured

Microsoft's latest Phi-4 excels at STEM, coding and multi-step reasoning. Strong math performance.

VRAM
10 GB
RAM
16 GB
Desktop tok/s
~35
Context
16k tokens
License
MIT
Available quants
Q4_K_MQ5_K_MQ8_0
reasoningcode generationmathdesktop
Qwen 2.5 0.5BQ8_0
0.5B params · 500 MB

Ultra-lightweight multilingual model for edge deployment. Fits on any device with a WebGPU runtime.

VRAM
800 MB
RAM
1 GB
Browser tok/s
~90
Desktop tok/s
~220
Mobile tok/s
~35
Context
33k tokens
License
Apache 2.0
Available quants
Q4_K_MQ8_0
tinyfastmultilingualbrowser-readyedge
Qwen 2.5 7BQ4_K_M
7B params · 4.7 GB

Strong multilingual and coding capabilities. Best run on desktop or mesh nodes.

VRAM
6 GB
RAM
8 GB
Browser tok/s
~12
Desktop tok/s
~48
Context
33k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_MQ8_0
multilingualcode generationdesktop
Qwen 2.5 Coder 7BQ4_K_M
7B params · 4.7 GB
Featured

Specialised coding variant of Qwen 2.5. Supports fill-in-middle (FIM) for IDE-style completions.

VRAM
6 GB
RAM
8 GB
Browser tok/s
~12
Desktop tok/s
~48
Context
33k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_MQ8_0
code generationfill-in-middledesktop
Qwen 2.5 72BQ4_K_M
72B params · 44.0 GB
Featured

Top-tier multilingual model with 128k context. Rivals GPT-4o on many benchmarks. Mesh-dispatched.

VRAM
49 GB
RAM
63 GB
Desktop tok/s
~12
Context
131k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_M
frontiermultilinguallong-contextmesh-only
Mistral 7B v0.3Q4_K_M
7B params · 4.1 GB

Reliable general-purpose model with strong instruction following. Wide hardware support.

VRAM
5 GB
RAM
8 GB
Browser tok/s
~13
Desktop tok/s
~52
Context
33k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_MQ8_0
generaldesktopinstruction
Mistral Nemo 12BQ4_K_M
12B params · 7.2 GB

Mistral × NVIDIA 12B with 128k context. Excellent multilingual performance and function-calling.

VRAM
9 GB
RAM
16 GB
Desktop tok/s
~38
Context
128k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_MQ8_0
long-contextmultilingualinstructiondesktop
Mixtral 8×7B MoEQ4_K_M
47B (MoE) params · 26.0 GB

Mixture-of-Experts with 8×7B experts, activating 13B per token. High quality at competitive speed.

VRAM
29 GB
RAM
47 GB
Desktop tok/s
~20
Context
33k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_M
moemesh-onlyhigh-qualitymultilingual
Gemma 2 2BQ4_K_M
2B params · 1.6 GB

Google's compact 2B model trained with knowledge distillation. Punches above its weight.

VRAM
2 GB
RAM
3 GB
Browser tok/s
~30
Desktop tok/s
~80
Mobile tok/s
~14
Context
8k tokens
License
Gemma
Available quants
Q4_K_MQ8_0
fastbrowser-readyinstructionedge
Gemma 2 9BQ4_K_M
9B params · 5.5 GB

Highly capable 9B from Google DeepMind. Competitive with models twice its size on many benchmarks.

VRAM
7 GB
RAM
12 GB
Desktop tok/s
~44
Context
8k tokens
License
Gemma
Available quants
Q4_K_MQ5_K_MQ8_0
high-qualitydesktopinstruction
Gemma 2 27BQ4_K_M
27B params · 16.5 GB

Google's most capable open Gemma model. Excellent reasoning and instruction following.

VRAM
20 GB
RAM
32 GB
Desktop tok/s
~22
Context
8k tokens
License
Gemma
Available quants
Q4_K_MQ5_K_M
high-qualitymesh-onlyreasoning
DeepSeek R1 7BQ4_K_M
7B params · 4.7 GB
Featured

Chain-of-thought reasoning specialist distilled from DeepSeek-R1. Excels at math and coding.

VRAM
6 GB
RAM
8 GB
Browser tok/s
~11
Desktop tok/s
~46
Context
128k tokens
License
MIT
Available quants
Q4_K_MQ5_K_MQ8_0
reasoninglong-contextmathcode generationdesktop
DeepSeek R1 14BQ4_K_M
14B params · 8.7 GB

Larger R1 distillation with stronger multi-step reasoning. Best for complex analysis tasks.

VRAM
11 GB
RAM
16 GB
Desktop tok/s
~30
Context
128k tokens
License
MIT
Available quants
Q4_K_MQ5_K_M
reasoninglong-contextmathcode generationmesh-only
DeepSeek Coder V2 16BQ4_K_M
16B (MoE) params · 9.8 GB

State-of-the-art coding MoE model. Supports 338 programming languages and 160k context.

VRAM
12 GB
RAM
16 GB
Desktop tok/s
~28
Context
164k tokens
License
DeepSeek
Available quants
Q4_K_MQ5_K_M
code generationfill-in-middlelong-contextmesh-only
Command R 35BQ4_K_M
35B params · 20.5 GB

Optimised for retrieval-augmented generation (RAG). Grounding, citations, and tool-use built in.

VRAM
23 GB
RAM
32 GB
Desktop tok/s
~18
Context
131k tokens
License
CC-BY-NC
Available quants
Q4_K_MQ5_K_M
raglong-contextmultilingualmesh-only
StarCoder2 3BQ8_0
3B params · 3.2 GB

Lightweight code generation with fill-in-middle. Trained on 600+ languages. Browser-eligible.

VRAM
4 GB
RAM
6 GB
Browser tok/s
~18
Desktop tok/s
~65
Mobile tok/s
~6
Context
16k tokens
License
BigCode OpenRAIL-M
Available quants
Q4_K_MQ8_0
code generationfill-in-middlebrowser-ready
StarCoder2 15BQ4_K_M
15B params · 9.1 GB

High-quality 15B code model with FIM. Best performance on competitive programming tasks.

VRAM
11 GB
RAM
16 GB
Desktop tok/s
~33
Context
16k tokens
License
BigCode OpenRAIL-M
Available quants
Q4_K_MQ5_K_MQ8_0
code generationfill-in-middledesktop
Nomic Embed v1.5Q8_0
137M params · 140 MB

High-quality text embeddings at near-zero cost. Runs in any browser, sub-5ms per batch.

VRAM
300 MB
RAM
512 MB
Browser tok/s
~180
Desktop tok/s
~500
Mobile tok/s
~60
Context
8k tokens
License
Apache 2.0
Available quants
Q8_0fp16
embeddingsfastbrowser-readyedge
BGE-M3Q8_0
570M params · 580 MB
Featured

State-of-the-art multilingual embedding model (100+ languages). Dense + sparse + ColBERT in one.

VRAM
900 MB
RAM
2 GB
Browser tok/s
~80
Desktop tok/s
~280
Mobile tok/s
~22
Context
8k tokens
License
MIT
Available quants
Q4_K_MQ8_0fp16
embeddingsmultilingualbrowser-readyrag
E5-Mistral 7BQ4_K_M
7B params · 4.3 GB

Instruction-tuned LLM repurposed for high-quality long-context embeddings. Best for RAG pipelines.

VRAM
5 GB
RAM
8 GB
Desktop tok/s
~40
Context
33k tokens
License
MIT
Available quants
Q4_K_MQ8_0
embeddingslong-contextragdesktop
LLaVA 1.6 7BQ4_K_M
7B params · 4.9 GB

Visual instruction-tuned model. Accepts images + text. Suitable for document and chart analysis.

VRAM
6 GB
RAM
10 GB
Desktop tok/s
~42
Context
4k tokens
License
Apache 2.0
Available quants
Q4_K_MQ5_K_M
visionmultimodaldesktop
Moondream 2Q8_0
1.86B params · 1.9 GB

Tiny vision-language model designed for edge deployment. Runs in-browser for image Q&A.

VRAM
3 GB
RAM
4 GB
Browser tok/s
~20
Desktop tok/s
~55
Mobile tok/s
~7
Context
2k tokens
License
Apache 2.0
Available quants
Q4_K_MQ8_0
visionmultimodalbrowser-readytiny