Resource hub

Models and providers

Pick and configure the model behind Atlas: frontier models, fast models, open weights, and fully local inference.

Models

Atlas with GPT-5.1 Codex: Plan First, Then Build in 2026

GPT-5.1 Codex in Atlas: a frontier model from OpenAI at $1.25 per Mtok input, $10 per Mtok output, built to execute an approved plan across a long tool chain.

Models

Atlas with Qwen3 235B-A22B (local via Ollama): Self-Hosted Flagship in 2026

Qwen3 235B-A22B (local via Ollama) in Atlas: roughly 140 GB at 4-bit, Free (self-hosted), 22B active of 235B total, and a brutal memory floor to plan for.

Models

Atlas with GPT-5.3 Codex Spark: Fast Interactive Edits in 2026

GPT-5.3 Codex Spark is the low latency Codex model: 128K context, 32K max output, $1.75 per Mtok input, $14 per Mtok output. Atlas setup and honest tradeoffs.

Models

Atlas with Qwen3.5 35B-A3B: The Cheapest Reasoning Tier of 2026

Qwen3.5 35B-A3B is the cheapest reasoning tier in the Qwen3.5 line at $0.25 per Mtok input and $2.00 per Mtok output, with 256K tokens (262,144) of context for Atlas.

Models

Atlas with Qwen2.5 32B Instruct: Cost, Context, and Setup in 2026

Run Atlas on Qwen2.5 32B Instruct in 2026. 128K tokens (131,072) of context, $0.70 per Mtok input, $2.80 per Mtok output, and an offline Ollama route.

Models

Atlas with GLM-4.7: The Value Pick of the GLM Line in 2026

GLM-4.7 drives Atlas at $0.60 per Mtok input and $2.20 per Mtok output on a 200K tokens (204,800) window, undercutting Kimi K2.6's $4.00 output rate.

Compare

Atlas vs Mistral Vibe for Code: Terminal AI Coding Agents in 2026

Compare Atlas and Mistral Vibe for Code in 2026. Atlas offers terminal-native TUI, explicit diffs, and BYO models. Mistral Vibe for Code provides a four-model stack, multi-platform access, and EU data sovereignty.

Models

Atlas with DeepSeek Coder 6.7B (Ollama): the thin-hardware fallback in 2026

DeepSeek Coder 6.7B (Ollama) in Atlas: a 3.8GB pull, roughly 6GB to serve, 16K tokens (16,384) of context, Free (self-hosted). Dated, but it starts fast.

Models

Atlas with Llama 3.2 1B (local via Ollama): the 2026 offline small_model slot

Llama 3.2 1B (local via Ollama) in Atlas: a 1.3GB CPU-friendly pull with a 128,000 token window, Free (self-hosted), best wired to small_model, not model.

Models

Atlas with GPT-5.3 Chat: The small_model Slot Pick for 2026

GPT-5.3 Chat is a non reasoning 128K model at $1.75 per Mtok input, $14 per Mtok output. Why it belongs in the Atlas small_model slot, not the main build loop.

Models

Atlas with Qwen Max: Setup, Cost, and the 32,768 Token Limit in 2026

Run Atlas on Qwen Max in 2026. Alibaba's original flagship costs $1.60 per Mtok input, $6.40 per Mtok output, with a 32K tokens (32,768) context window.

Models

Atlas with GLM-4.5: The MIT-Licensed Open Weights Baseline in 2026

GLM-4.5 runs Atlas at $0.60 per Mtok input and $2.20 per Mtok output on a 128K tokens (131,072) window, under an MIT license that permits commercial self-hosting.

Models

Atlas with GPT-5.2 Pro: The Deep Reasoning Tier in 2026

Running Atlas on GPT-5.2 Pro in 2026: a 400K token context, $21 per Mtok input, $168 per Mtok output, and the one shot questions that justify the price.

Models

Atlas with AllenAI Olmo 3 32B Think: the fully open reasoning model in 2026

Run Atlas on AllenAI Olmo 3 32B Think in 2026: 65,536 token context, OpenRouter $0.15/$0.50 per Mtok, and the only weights, data, and training code you can audit.

Models

Atlas with GLM-5: Z.ai's Frontier Tier for Coding Agents in 2026

GLM-5 runs Atlas at $1.00 per Mtok input and $3.20 per Mtok output on a 200K tokens (204,800) window. Z.ai's February 2026 jump past the 4.x line.

Models

Atlas with Gemini 2.5 Flash-Lite: A 1M Context Helper Model for $0.1 per Mtok in 2026

Gemini 2.5 Flash-Lite in Atlas: $0.1 per Mtok input, $0.4 per Mtok output, a 1,048,576 token context, and reasoning enabled. Google's cheapest million token model.

Models

Atlas with MiniMax-M2.5-highspeed in 2026: Paying 2x for Latency

MiniMax-M2.5-highspeed runs Atlas at $0.60 per Mtok input and $2.40 per Mtok output, exactly double base M2.5, for identical weights and a 204,800 token context.

Models

Atlas with Claude Opus 4.5: The Price Break Opus in 2026

Claude Opus 4.5 runs Atlas on a 200K context at $5 per Mtok input, $25 per Mtok output, down from $15/$75 on Opus 4.1. Setup, extended thinking, and tradeoffs.

Models

Atlas with Grok 4.20 (Non-Reasoning) in 2026: Fast, Predictable Edit Passes

Grok 4.20 (Non-Reasoning) bills no thinking tokens, so cost per turn is fully predictable at $1.25 / $2.5 per Mtok, with the same 1,000,000 token context as the reasoning variant.

Models

Atlas with Claude Opus 4.6: Setup, Cost, and Tradeoffs in 2026

Claude Opus 4.6 gives Atlas a 1M token context at $5 per Mtok input, $25 per Mtok output. Setup steps, cost math, and when to pick a newer Opus instead.

Models

Atlas with GPT-5.6 Luna: Best Price Per Context in 2026

GPT-5.6 Luna keeps the full 1,050,000 token window at $1 / $6 per Mtok, one fifth of GPT-5.6. The best price-per-context ratio for agentic coding in Atlas in 2026.

Models

Atlas with Qwen3.6 Max Preview: The Top 3.6 Tier, Reviewed in 2026

Qwen3.6 Max Preview is Alibaba's top Qwen3.6 tier at $1.30 per Mtok input and $7.80 per Mtok output, with 256K tokens (262,144) of context. What it is worth inside Atlas.

Models

Atlas with GPT-5.5 Pro: One Hard Question at a Time in 2026

GPT-5.5 Pro is OpenAI's April 2026 maximum-effort reasoning tier at $30 / $180 per Mtok on a 1,050,000 token window. Invoke it deliberately in Atlas, then switch back.

Models

Atlas with Qwen3 14B: The Middle Dense Tier Worth Pinning in 2026

Qwen3 14B in Atlas for 2026: reasoning-enabled dense 14B at $0.35 per Mtok input and $1.40 per Mtok output, 128K tokens (131,072), roughly 9 GB quantized.

Models

Atlas with Llama 3.1 405B: The 243GB Landmark in 2026

Llama 3.1 405B in Atlas for 2026: 405 billion openly released parameters, a 128,000 token window, Free (self-hosted), and a 243GB download that decides everything.

Models

Atlas with Qwen3.7 Plus: The Newest Million-Token Qwen Tier, June 2026

Qwen3.7 Plus is the value play in Alibaba's 3.7 line: 1M tokens (1,000,000) of context at $0.50 per Mtok input and $3.00 per Mtok output, one fifth the input price of Qwen3.7 Max.

Models

Atlas with Gemini 3 Flash (2026): The $0.5 Default for All Day Agent Loops

Gemini 3 Flash gives Atlas a 1,048,576 token context and 65,536 token output at $0.5 per Mtok input and $3 per Mtok output. The economical default for agent loops.

Models

Atlas with GPT-5.6 Sol: The High Effort GPT-5.6 Variant in 2026

GPT-5.6 Sol tops the 5.6 line at $5 per Mtok input, $30 per Mtok output on a 1,050,000 token window. Atlas setup, small_model routing, and when Sol is overkill.

Models

Atlas with Qwen3 14B (Ollama): the local planning model for 2026

Qwen3 14B (Ollama) in Atlas: dense 9.3GB weights, roughly 11GB to serve, 40K tokens (40,960) of context, Free (self-hosted), better at architecture than raw diffs.

Models

Atlas with Mistral 7B v0.3 (Ollama): The Predictable Local Baseline in 2026

Run Atlas on Mistral 7B v0.3 (Ollama): a 4.4GB Apache 2.0 model with a 32K context, free self-hosted. Setup, tradeoffs, and when to pick a stronger coder.

Models

Atlas with GPT-5.2 Codex: Agentic Coding at $1.75 per Mtok in 2026

GPT-5.2 Codex in Atlas: 400K context, $1.75 per Mtok input and $14 per Mtok output, with Codex post training for long running software engineering work.

Models

Atlas with DeepSeek-R1 32B Distill (Ollama): Local Reasoning on a 24GB Card in 2026

DeepSeek-R1 32B Distill (Ollama) is 20GB of weights with a 128K context, the strongest R1 distill that fits a 24GB consumer GPU. Atlas setup and tradeoffs for 2026.

Models

Atlas with Grok 4.5: A Symmetric 500K Context and Output Budget (2026)

Grok 4.5 drives Atlas at $2 / $6 per Mtok with a 500K token context and a matching 500,000 max output tokens. Atlas routes xAI through the Responses API.

Models

Atlas with GPT-5.6: OpenAI's 2026 Flagship in the Terminal

GPT-5.6 gives Atlas a 1,050,000 token context at $5 / $30 per Mtok. Atlas routes OpenAI through the Responses API so reasoning persists across tool calls in 2026.

Models

Atlas with GPT-4.1 (2026): A Million Token Window at $2 In, $8 Out

GPT-4.1 gives Atlas a 1,047,576 token context at $2 per Mtok input and $8 per Mtok output. Fast, non reasoning, with a 32,768 token output cap. Setup and tradeoffs.

Models

Atlas with Qwen2.5-Coder 32B (Ollama): the air-gapped ceiling in 2026

Qwen2.5-Coder 32B (Ollama) in Atlas: 20GB of weights, roughly 22GB to serve on a 24GB card, 32K tokens (32,768) of context, Free (self-hosted), fully air-gapped.

Models

Atlas with GPT-5.2: The 400K General Purpose GPT in 2026

GPT-5.2 runs Atlas on a 400K context with 128K max output at $1.75 per Mtok input, $14 per Mtok output. Setup, the price rise from GPT-5.1, and the Codex tradeoff.

Models

Atlas with Kimi K2.5: The Cheap Reasoning Default for 2026

Kimi K2.5 gives Atlas a current-generation reasoning model at $0.60 per Mtok input and $3.00 per Mtok output, with a full 256K tokens (262,144) context window.

Models

Atlas with GPT-4.1 mini (2026): The Cheap Slot That Still Holds a Million Tokens

GPT-4.1 mini gives Atlas a 1,047,576 token context at $0.40 per Mtok input and $1.60 per Mtok output. The right small_model when your cheap slot needs a huge window.

Models

Atlas with GPT-OSS 20B (hosted): the cheap slot that can still think in 2026

Run Atlas on GPT-OSS 20B (hosted) in 2026: $0.03/$0.14 per Mtok on DeepInfra, a 131,072 token context, reasoning on, and $0.00/$0.00 locally in LM Studio.

Models

Atlas with Nemotron 70B (Ollama): Instruction Adherence on a Workstation in 2026

Run Atlas on Nemotron 70B (Ollama): NVIDIA's 43GB reward-tuned Llama 3.1 70B with a 128K context, free self-hosted. Setup, memory math, and honest tradeoffs.

Models

Atlas with Databricks Foundation Model APIs in 2026: Frontier Models Inside the Lakehouse

Atlas with Databricks Foundation Model APIs in 2026: Claude Opus 4.7 at a 1,000,000 token context, GPT-5.5 at 1,050,000, governed by the same DATABRICKS_TOKEN.

Models

Atlas with DeepSeek-R1 1.5B Distill (Ollama): The 1.1GB Reasoning Slot in 2026

DeepSeek-R1 1.5B Distill (Ollama) is a 1.1GB reasoning model with a 128K context that runs on CPU. Use it as the Atlas small_model in 2026. Free (self-hosted).

Models

Atlas with StarCoder2 15B (Ollama): The Provenance Choice in 2026

StarCoder2 15B (Ollama) is BigCode's 9.1GB code model with a transparent training corpus and a 16K context, Free (self-hosted). Atlas setup and tradeoffs for 2026.

Models

Atlas with OpenAI o3-pro: The $20 / $80 Reasoning Escape Hatch (2026)

OpenAI o3-pro in Atlas costs $20 per Mtok input and $80 per Mtok output on a 200K context. A one shot escape hatch for hard problems, not an interactive default.

Models

Atlas with GPT-5 Chat: The Only Chat Snapshot That Keeps 400K in 2026

GPT-5 Chat in Atlas: gpt-5-chat-latest is the one chat-latest snapshot that retains a full 400K context and 128K output, at $1.25 per Mtok input.

Models

Atlas with NVIDIA Nemotron Nano 9B v2 in 2026

Nemotron Nano 9B v2 in Atlas, 2026: a dense 9B reasoning model at $0.06/$0.23 per Mtok on Vercel AI Gateway and Amazon Bedrock, free on the NVIDIA NIM tier.

Models

Atlas with OpenCoder 8B (Ollama): the Auditable Local Coder in 2026

OpenCoder 8B (Ollama) runs Atlas on a fully open code LLM: open data, open recipe, 4.7GB, 8K tokens (8,192) of context, Free (self-hosted). Setup and tradeoffs.

Models

Atlas with IBM Granite 4 Small-H (Ollama): a 1M-Token Local Window in 2026

IBM Granite 4 Small-H (Ollama) is a hybrid Mamba model with 1M tokens (1,048,576) of context from a 19GB download, Free (self-hosted). Atlas setup and memory notes.

Models

Atlas with Together AI (gateway) in 2026: Open-Weights Models at Scale

Together AI (gateway) drives Atlas with the broadest open-weights catalog: Qwen3.7 Max at $1.25 / $3.75 per Mtok, up to 1M context, US-hosted inference.

Models

Atlas with GPT-4o in 2026: A 128K Legacy Model with Dated Snapshots

GPT-4o runs in Atlas at $2.50 per Mtok input and $10 per Mtok output on a 128K context with a 16,384 output cap. Best for quick lookups and reproducible baselines.

Models

Atlas with Llama 3.3 70B Instruct (Meta Llama API) in 2026

Llama 3.3 70B Instruct (Meta Llama API) in Atlas for 2026: 128,000 tokens of context, an OpenAI-compatible endpoint, and a 4,096 token output ceiling to plan around.

Models

Atlas with Claude Sonnet 4.6: The Build Agent Default in 2026

Claude Sonnet 4.6 drives Atlas at $3 per Mtok input, $15 per Mtok output on a 1M token window with 128K max output. Setup, cost math, and when Sonnet 5 wins.

Models

Atlas with Azure OpenAI (gateway) in 2026: The GPT-5 Codex Line Under Enterprise IAM

Azure OpenAI (gateway) runs Atlas on the full GPT-5 lineup including gpt-5.3-codex, with up to 1.05M context and passthrough Azure pricing in your own region.

Models

Atlas with Gemma 3 12B Instruct: The Mid-Size Bedrock Gemma for 2026

Gemma 3 12B Instruct in Atlas via Amazon Bedrock: $0.05 per Mtok input, $0.10 per Mtok output, a 131,072 token context, and 12B dense weights you can self-host.

Models

Atlas with GLM-4.7-FlashX: The Cheapest Paid Slot in 2026

GLM-4.7-FlashX runs Atlas at $0.07 per Mtok input and $0.40 per Mtok output on a 200K tokens (200,000) context, with reasoning enabled and a 131,072 output cap.

Models

Atlas with GPT-5 Pro: The 272,000 Token Output Ceiling in 2026

GPT-5 Pro in Atlas: the only OpenAI model with a 272,000 token max output, priced at $15 per Mtok input, $120 per Mtok output on a 400K tokens window.

Models

Atlas with Devstral Small 2: A Free Coding Agent Model in 2026

Devstral Small 2 in Atlas: listed at $0 / $0 per Mtok on Mistral's labs endpoint with a 256K window, or run 24B locally at roughly 14GB via ollama pull devstral.

Models

Atlas with Llama 3.2 3B (local via Ollama): A 2.0GB Offline small_model for 2026

Llama 3.2 3B (local via Ollama) in Atlas for 2026: a 2.0GB pull, Free (self-hosted), 128,000 tokens of context, and the offline small_model that never touches a network.

Models

Atlas with DeepSeek Coder V2 16B Lite (Ollama): The 160K Local Agent Model in 2026

DeepSeek Coder V2 16B Lite (Ollama) gives Atlas a 160K token context from an 8.9GB download, Free (self-hosted). Setup, MoE speed, and the KV cache catch, for 2026.

Models

Atlas with Qwen2.5-Coder 7B (local via Ollama): the Laptop Setup in 2026

Qwen2.5-Coder 7B runs Atlas on a laptop with no discrete GPU: about 5GB at 4-bit, a 32,768 token context, and free self-hosted. Setup, limits, and when to upgrade.

Models

Atlas with Amazon Nova Pro: Bedrock Billing and a 300K Window (2026)

Amazon Nova Pro runs Atlas at $0.80 / $3.20 per Mtok on a 300,000 token window, billed through your existing AWS account with no new vendor contract.

Models

Atlas with Ollama Cloud in 2026: Hosted Ollama Tags for a Terminal Coding Agent

How to run Atlas on Ollama Cloud in 2026: same local tags like qwen3-coder:480b on hosted GPUs, up to 1,048,576 tokens of context, no published per-token price.

Models

Atlas with Qwen3 32B: The Largest Dense Qwen3 in 2026

Qwen3 32B in Atlas: 16,384 max output tokens, hybrid thinking, 128K tokens (131,072) of context, $0.70 per Mtok input and $2.80 per Mtok output in 2026.

Models

Atlas with MiniMax-M2.1 in 2026: A Free Upgrade Over M2

MiniMax-M2.1 runs Atlas at $0.30 per Mtok input and $1.20 per Mtok output with a 204,800 token context and 131,072 max output. Setup, tradeoffs, and when to move on.

Models

Atlas with Amazon Bedrock (gateway) in 2026: One IAM Boundary, Every Model

Amazon Bedrock (gateway) drives Atlas behind AWS IAM: Nova Lite at $0.06 / $0.24 per Mtok, Llama 4 Scout with a 3,500,000 token context, full AWS auth chain.

Models

Atlas with Qwen3-Coder 30B-A3B Instruct: The Default Open Agentic Coder in 2026

Qwen3-Coder 30B-A3B Instruct in Atlas: 256K tokens (262,144), $0.45 per Mtok input, $2.25 per Mtok output, and 3.3B active parameters out of 30B total.

Models

Atlas with Claude Sonnet 5: The 2026 Everyday Driver

Claude Sonnet 5 gives Atlas a 1M token window at $2 / $10 per Mtok, 60 percent under Opus 4.8 input pricing. Why it is the default for long agentic sessions in 2026.

Models

Atlas with Llama 3.3 70B (local via Ollama) in 2026: Dense 70B on Your Own Hardware

Llama 3.3 70B (local via Ollama) drives Atlas at 128K tokens (131,072) and Free (self-hosted) pricing, with a $0.59 / $0.79 per Mtok Groq fallback on identical weights.

Models

Atlas with GPT-5.3 Codex: Code-Specialized Reasoning in 2026

GPT-5.3 Codex is OpenAI's February 2026 code-specialized reasoning model, $1.75 / $14 per Mtok on a 400K window. Built for the long agentic loops Atlas runs.

Models

Atlas with Kimi K2 Thinking Turbo: The 2026 Reasoning Speed Tier

Kimi K2 Thinking Turbo gives Atlas priority serving on a reasoning model at $1.15 per Mtok input and $8.00 per Mtok output, on a 256K tokens (262,144) window.

Models

Atlas with Mistral Medium 3.5: The EU-Hosted Frontier Model (2026)

Mistral Medium 3.5 drives Atlas at $1.50 / $7.50 per Mtok on a 262,144 token window with a matching 262,144 output limit. The EU-hosted option for data residency.

Models

Atlas with GPT-5.6 Terra: The Mid Tier GPT-5.6 Pick for 2026

GPT-5.6 Terra drives Atlas on a 1,050,000 token context at $2.50 per Mtok input, $15 per Mtok output. Setup, cost math, and when Sol or Luna is the better pick.

Models

Atlas with GPT-5.5: The 1.05M Token Jump, Reviewed for 2026

GPT-5.5 took the GPT-5 line to a 1,050,000 token context at $5 per Mtok input, $30 per Mtok output. Atlas setup, the price rise from $1.75, and when GPT-5.6 wins.

Models

Atlas with Claude Opus 4.7 in 2026: 1M Context at $5 / $25 per Mtok

Claude Opus 4.7 drives Atlas with a 1M tokens (1,000,000) window, 128K max output, and $5 per Mtok input, $25 per Mtok output. Setup, tradeoffs, and when to move on.

Models

Atlas with Mixtral 8x22B (local via Ollama): The 80GB Question in 2026

Mixtral 8x22B (local via Ollama) in Atlas for 2026: an 80GB pull, Free (self-hosted), a 64,000 token window, and sparse routing across 8 experts of 22B.

Models

Atlas with Qwen3-Coder 30B (local via Ollama): the Default Local Setup in 2026

Qwen3-Coder 30B is Atlas's default local coding model: a 19GB Ollama download, 256K context, 3.3B active parameters, and free self-hosted. Setup and real tradeoffs.

Models

Atlas with Fireworks AI (gateway) in 2026: Buying Latency with Money

Fireworks AI (gateway) serves Atlas open models with fast-router tiers: DeepSeek V4 Flash at $0.14 / $0.28 per Mtok on a 1,000,000 token context.

Models

Atlas with Grok 4.3: The Cheapest 1M Context Reasoning Model (2026)

Grok 4.3 gives Atlas a 1M token context at $1.25 / $2.50 per Mtok, the cheapest 1M reasoning model from a US lab. Output caps at 30,000 tokens, so plan around it.

Models

Atlas with Kimi K2.7 Code: Open Weights at Trillion-Parameter Scale (2026)

Kimi K2.7 Code drives Atlas at $0.95 / $4 per Mtok on a 262,144 token window. Open weights, 1T total parameters, served by four independent providers.

Models

Atlas with Codestral: Fast Fill-in-the-Middle Editing in the Terminal (2026)

Codestral runs in Atlas at $0.30 / $0.90 per Mtok on a 256K token window. Fast single-file edits, but a 4,096 token output ceiling blocks large refactors.

Models

Atlas with GPT-5 Codex: The First Codex Model of the GPT-5 Family in 2026

GPT-5 Codex in Atlas: September 2025 brought coding post training at zero premium, 400K context, 128K output, and $1.25 per Mtok input with $10 per Mtok output.

Models

Atlas with DeepSeek-V3.1 671B (Ollama): Self-Hosting Frontier Open Weights in 2026

DeepSeek-V3.1 671B (Ollama) is a 404GB MoE with a 160K context and hybrid thinking modes. Run Atlas on frontier open weights behind your own firewall in 2026.

Models

Atlas with IBM Granite Code 8B (Ollama): 125K Context on 4.6GB in 2026

IBM Granite Code 8B (Ollama) gives Atlas a 125K tokens window from a 4.6GB download, Free (self-hosted), with enterprise licensing. Setup, tags, and tradeoffs.

Models

Atlas with GPT-5 Mini: Full 400K Context at One Fifth the Price in 2026

GPT-5 Mini in Atlas: $0.25 per Mtok input and $2 per Mtok output, 5x cheaper than GPT-5 on both sides, with no reduction to the 400K context window.

Models

Atlas with Mistral Medium 3.1 (2508): Setup, Cost, and Fit in 2026

Run Atlas on Mistral Medium 3.1 (2508): a 262,144 token window at $0.40 / 1M input tokens and $2.00 / 1M output tokens. Setup steps, costs, and honest tradeoffs.

Models

Atlas with Mistral 7B: Cost, Context, and Real Limits in 2026

Running Atlas on Mistral 7B in 2026: an 8,000 token window at $0.25 / 1M input tokens. Great for smoke-testing a provider block, wrong for agentic coding.

Models

Atlas with OpenAI o3-mini in 2026: Still Worth Pinning?

OpenAI o3-mini runs in Atlas at $1.10 per Mtok input and $4.40 per Mtok output on a 200K context, but o4-mini costs exactly the same and is newer. Here is the call.

Models

Atlas with OpenAI o3: Cheap Deep Reasoning in the Terminal (2026)

Run Atlas on OpenAI o3 in 2026. A 200K token context reasoning model at $2 per Mtok input and $8 per Mtok output, with real setup steps and honest tradeoffs.

Models

Atlas with DeepSeek Chat: 384,000 Token Output at $0.28 per Mtok in 2026

Run Atlas on DeepSeek Chat in 2026. DeepSeek's non-reasoning endpoint gives 1M tokens (1,000,000) of context at $0.14 per Mtok input, $0.28 per Mtok output.

Models

Atlas with Phi-4 (local via Ollama): The 16K Context Tradeoff in 2026

Phi-4 (local via Ollama) drives Atlas at Free (self-hosted) pricing with a 16K tokens (16,384) window. Setup, the 14B reasoning case, and where it breaks.

Models

Atlas with Gemma 3 27B (local via Ollama): Setup, Cost, and Tradeoffs in 2026

Run Atlas on Gemma 3 27B (local via Ollama) in 2026: a 131,072 token context, Free (self-hosted) pricing, single-GPU inference, and the honest tradeoffs.

Models

Atlas with Gemini 3.1 Pro: Setup, Cost, and Tradeoffs in 2026

Run Atlas on Gemini 3.1 Pro in 2026: a 1,048,576 token window at $2 / $12 per Mtok, with real setup steps, the 65,536 output ceiling, and when to switch models.

Models

Atlas with Amazon Nova 2 Lite: Cost, Context, and Setup in 2026

Amazon Nova 2 Lite drives Atlas from inside your AWS account at $0.33 / $2.75 per Mtok with a 128K context. Setup, real tradeoffs, and when to pick another model.

Models

Atlas with Gemini 3.5 Flash: The May 2026 Latency Tier That Keeps 1M Context

Gemini 3.5 Flash in Atlas: the May 2026 release keeps the full 1,048,576 token window at $1.50 / $9 per Mtok, with reasoning and tool calling enabled.

Models

Atlas with Qwen2.5-Coder 1.5B (Ollama): the 986MB small_model slot in 2026

Qwen2.5-Coder 1.5B (Ollama) in Atlas: a 986MB Q4_K_M pull with 32K tokens (32,768) of context, Free (self-hosted), sized for titles and summaries, not refactors.

Models

Atlas with Claude Opus 4.8: The 2026 Model Guide

Claude Opus 4.8 in Atlas: a 1M token context window at $5 / $25 per Mtok with 128K output. When to pin Anthropic's top coding model in 2026, and when not to.

Models

Atlas with DeepSeek V4 Flash: The Cheapest 1M Context Reasoning Model in 2026

DeepSeek V4 Flash in Atlas: $0.14 / $0.28 per Mtok on a 1M window with 384,000 output tokens, or $0.09 / $0.18 via DeepInfra. Setup, small_model wiring, tradeoffs.

Models

Atlas with Mistral Small 3.2 (2506): The Cheap Slot Done Right in 2026

Mistral Small 3.2 (2506) gives Atlas a 128,000 token window at $0.10 / 1M input tokens and $0.30 / 1M output tokens. Setup, the 16,384 token output cap, tradeoffs.

Models

Atlas with Gemini 2.5 Pro: The Cheap 1M Context Default in 2026

Gemini 2.5 Pro in Atlas: a stable GA id with a 1,048,576 token window at $1.25 per Mtok input, 37 percent cheaper to read with than Gemini 3 Pro.

Models

Atlas with GPT-5.1 Codex mini: Wide Subagent Fan Out in 2026

GPT-5.1 Codex mini in Atlas: an OpenAI fast tier at $0.25 per Mtok input, $2 per Mtok output, built for parallel subagents, bulk sweeps, and cheap summaries.

Models

Atlas with Magistral Small (local via Ollama): Private Reasoning in 2026

Magistral Small (local via Ollama) in Atlas for 2026: a 14GB pull, Free (self-hosted), reasoning traces generated on your own GPU, and an honest 40,000 token cap.

Models

Atlas with Command R: The $0.15 per Mtok Background Model for 2026

Command R runs Atlas background work at $0.15 per Mtok input and $0.6 per Mtok output, with a 128,000 token context and a 4,000 token output cap on patches.

Models

Atlas with Qwen3 8B: The Cheap Thinking Model for small_model in 2026

Qwen3 8B in Atlas for 2026: hybrid thinking at $0.18 per Mtok input and $0.70 per Mtok output, 128K tokens (131,072), and a natural small_model slot fit.

Models

Atlas with OpenAI o1 in 2026: The Original Reasoning Model, Priced Like It

OpenAI o1 runs in Atlas at $15 per Mtok input and $60 per Mtok output on a 200K context. Historically important, now dominated by o3 on both price and capability.

Models

Atlas with QwQ Plus: Alibaba's Reasoning Line in the Plan Agent, 2026

Run Atlas on QwQ Plus in 2026. Alibaba's dedicated reasoning tier costs $0.80 per Mtok input, $2.40 per Mtok output, with a 128K tokens (131,072) context.

Models

Atlas with Gemini Flash-Lite Latest: A Set-and-Forget small_model for 2026

gemini-flash-lite-latest in Atlas: a rolling alias for Google's cheapest reasoning-capable Lite tier at $0.1 per Mtok input and a 1,048,576 token context.

Models

Atlas with Vercel AI Gateway in 2026: 310 Models Behind One AI_GATEWAY_API_KEY

Atlas with Vercel AI Gateway in 2026: roughly 310 models, Grok 4.20 Reasoning at a 2,000,000 token context for $1.25/$2.50 per Mtok, one AI_GATEWAY_API_KEY.

Models

Atlas with GLM-5.1: Reasoning, Cost, and Context in 2026

GLM-5.1 from Z.ai drives Atlas with a 200,000 token context at $1.40 per Mtok input and $4.40 per Mtok output. Setup, honest tradeoffs, and how it compares to GLM-5.2.

Models

Atlas with Qwen3-Coder 480B (Ollama): the self-hosted ceiling in 2026

Qwen3-Coder 480B (Ollama) in Atlas: 290GB of weights, roughly 292GB to serve, 256K tokens (262,144) of context. Free (self-hosted), but the hardware is not.

Models

Atlas with IBM Granite 4.1 8B in 2026

IBM Granite 4.1 8B in Atlas, 2026: a dense 8B model at $0.05/$0.10 per Mtok with a 131,072 token context where max output equals the full window.

Models

Atlas with Snowflake Cortex in 2026: Claude Opus 4.8 Inside Your Snowflake Boundary

Atlas with Snowflake Cortex in 2026: Claude Opus 4.8 and Claude Fable 5 at a 1,000,000 token context, governed by Snowflake RBAC and billed in Snowflake credits.

Models

Atlas with Code Llama 34B (Ollama): The Practical Top of the Line in 2026

Code Llama 34B (Ollama) is 19GB, roughly 21GB to serve, with a 16K context and Free (self-hosted) pricing. Atlas setup, why 34B is the last sane size, for 2026.

Models

Atlas with Llama 3.3 8B Instruct (Meta Llama API): The small_model Slot in 2026

Llama 3.3 8B Instruct (Meta Llama API) in Atlas for 2026: 128,000 tokens of context in an 8B-class model, a 4,096 token output cap, and why it belongs in small_model.

Models

Atlas with Amazon Nova Lite in 2026: 300K Context on Your AWS Bill

Amazon Nova Lite drives Atlas at $0.06 per Mtok input and $0.24 per Mtok output with a 300K token context, billed through IAM with no new vendor API key to manage.

Models

Atlas with Mistral Small 3.2 (local via Ollama) in 2026

Run Atlas on Mistral Small 3.2 (local via Ollama) in 2026: a 15GB pull, Free (self-hosted), 128,000 tokens of context, and function-calling tuned 2506 weights.

Models

Atlas with Code Llama 7B (Ollama): A 3.8GB Local Coder in 2026

Code Llama 7B (Ollama) is Meta's 2023 code model at 3.8GB with a 16K context, Free (self-hosted). Atlas setup, the :code and :python tags, and 2026 tradeoffs.

Models

Atlas with Ministral 3B: The $0.04 Housekeeping Model in 2026

Ministral 3B is the cheapest model Mistral sells: $0.04 / 1M input tokens and $0.04 / 1M output tokens across 128,000 tokens. Atlas small_model setup and limits.

Models

Atlas with Hugging Face Inference in 2026: 51 Models Behind One HF_TOKEN

Atlas with Hugging Face Inference in 2026: 51 routed models behind one HF_TOKEN, GLM-4.7-Flash free at $0/$0 per Mtok, and the router margin on popular rows.

Models

Atlas with GPT-5.4 Pro: One Shot Hard Problems in 2026

GPT-5.4 Pro is the max reasoning tier at $30 per Mtok input, $180 per Mtok output on a 1,050,000 token window. Why Atlas users switch to it instead of pinning it.

Models

Atlas with QwQ 32B (Ollama): a free local reasoning model for the plan agent in 2026

QwQ 32B (Ollama) in Atlas: Qwen's dedicated reasoning model at 20GB, 40K tokens (40,960) of context, Free (self-hosted). Let QwQ plan, then hand off to a coder.

Models

Atlas with Google Vertex AI (gateway) in 2026: Gemini and Claude Under One GCP Project

Google Vertex AI (gateway) runs Atlas on Gemini 3.1 Pro at $2 / $12 per Mtok with a 1M context, plus Claude through Atlas's google-vertex-anthropic route.

Models

Atlas with Qwen3.6 35B-A3B: The $0.248 per Mtok Workhorse of 2026

Qwen3.6 35B-A3B is the cheapest reasoning model in the Qwen3.6 line at $0.248 per Mtok input and $1.485 per Mtok output, with a full 256K tokens (262,144) context for Atlas.

Models

Atlas with Qwen3 4B (Ollama): 256K of context on a 2.5GB pull in 2026

Qwen3 4B (Ollama) in Atlas: a 2.5GB download carrying 256K tokens (262,144) of context, Free (self-hosted), with separate 2507 instruct and thinking tags.

Models

Atlas with GPT-5 Nano: The Cheapest Model in the OpenAI Registry in 2026

GPT-5 Nano in Atlas: $0.05 per Mtok input and $0.40 per Mtok output, the cheapest model in the OpenAI registry, and the right pick for the small_model slot.

Models

Atlas with Qwen2.5 72B Instruct: The Flagship Dense Qwen in 2026

Qwen2.5 72B Instruct in Atlas: 128K tokens (131,072), $1.40 per Mtok input, $5.60 per Mtok output, openly published weights you can serve on your own vLLM.

Models

Atlas with DeepInfra: The Cheapest Open-Weights Host for an Agent Loop in 2026

Run Atlas on DeepInfra: GPT OSS 120B at $0.037/$0.17 per Mtok, DeepSeek V4 Flash at a 1,048,576 token window for $0.09 input. Setup, limits, and cost math.

Models

Atlas with Grok 4.20 Multi-Agent in 2026: Two Levels of Fan-Out

Grok 4.20 Multi-Agent orchestrates internal agents behind one model id, stacking with Atlas's own parallel subagents. 1,000,000 token context at $1.25 / $2.5 per Mtok.

Models

Atlas with Magistral Small: Open Reasoning for the Plan Agent in 2026

Magistral Small is Mistral's first open reasoning model: 128,000 tokens at $0.50 / 1M input tokens and $1.50 / 1M output tokens. Atlas setup, costs, and tradeoffs.

Models

Atlas with Kimi K2 Thinking: Open-Weights Reasoning at $0.60 per Mtok (2026)

Kimi K2 Thinking gives Atlas reasoning at $0.60 / $2.50 per Mtok on a 262,144 token window. Open weights, and it sustains the long tool-calling chains agents need.

Models

Atlas with Kimi K2 0905: 262,144 Tokens In and Out, 2026

Run Atlas on Kimi K2 0905 in 2026. Moonshot's September refresh gives 256K tokens (262,144) context and output at $0.60 per Mtok in, $2.50 per Mtok out.

Models

Atlas with Qwen Flash: A Cheap Tier That Still Writes Real Diffs in 2026

Run Atlas on Qwen Flash in 2026. Alibaba's fast tier pairs 1M tokens (1,000,000) of context with a 32,768 token output at $0.05 per Mtok in, $0.40 per Mtok out.

Models

Atlas with Code Llama (local via Ollama): a Fill-in-the-Middle Baseline in 2026

Code Llama runs free and local via Ollama at 7b, 13b, 34b, and 70b with a 16,384 token context. Strong at infilling, but too small and too old for Atlas agent work.

Models

Atlas with Qwen3.6 Flash: A $0.1875 Speed Tier With Coding Lineage in 2026

Qwen3.6 Flash in Atlas: Alibaba's April 2026 speed tier at $0.1875 / $1.125 per Mtok with a 1M window, wired to small_model in atlas.json. Setup and tradeoffs.

Models

Atlas with Qwen3 Coder Plus: Agentic Coding on a 1M Window in 2026

Qwen3 Coder Plus in Atlas: a 1,048,576 token window at $1 / $5 per Mtok, post-trained for agentic coding, with ollama pull qwen3-coder:30b as the local counterpart.

Models

Atlas with Llama 4 Scout: the 3.5M Token Context Model in 2026

Llama 4 Scout gives Atlas a 3.5M token context on Bedrock at $0.17 / $0.66 per Mtok, or $0.10 / $0.30 on DeepInfra. Setup, the portability trap, and when to switch.

Models

Atlas with GPT-5.4: The Balanced 2026 Default

GPT-5.4 brings a 1,050,000 token window to Atlas at $2.50 / $15 per Mtok, half the input cost of GPT-5.6. The balanced day-to-day default for 2026 sessions.

Models

Atlas with Gemma 3 12B (Ollama): 128K Context and Screenshots in 2026

Gemma 3 12B (Ollama) gives Atlas a 128K tokens (131,072) window and multimodal image input from an 8.1GB download, Free (self-hosted). Setup, sizing, and limits.

Models

Atlas with DeepSeek Coder 33B (Ollama): Local Setup and Tradeoffs in 2026

DeepSeek Coder 33B (Ollama) drives Atlas locally from 19GB of weights with a 16K token context and no API bill. Setup, VRAM budget, and honest tradeoffs for 2026.

Models

Atlas with AI21 Jamba Large 1.7 in 2026

AI21 Jamba Large 1.7 in Atlas, 2026: a hybrid SSM-Transformer with 256,000 tokens of context at $2.00/$8.00 per Mtok, capped at 4,096 tokens of output.

Models

Atlas with Codestral 22B (Ollama): 32K Context and a License Gate in 2026

Codestral 22B (Ollama) is Mistral's 13GB code model with a 32K context, fluent across many languages. Free (self-hosted), non-commercial license. Atlas setup for 2026.

Models

Atlas with Upstage Solar Pro 3 in 2026

Solar Pro 3 in Atlas, 2026: symmetric $0.25/$0.25 per Mtok on the Upstage API at 131,072 tokens, versus $0.15/$0.60 and 128,000 tokens on OpenRouter.

Models

Atlas with Mistral Medium 3 (2505) in 2026: Symmetric Limits, $0.40 In

Mistral Medium 3 (2505) runs Atlas with symmetric 131,072 token context and output at $0.40 / 1M input tokens and $2.00 / 1M output tokens. Setup, limits, successors.

Models

Atlas with Magistral Medium: Multi-Hop Root-Cause Debugging in 2026

Magistral Medium reasons across a 128,000 token window at $2.00 / 1M input tokens and $5.00 / 1M output tokens. Atlas setup, the 16,384 token output cap, tradeoffs.

Models

Atlas with Cerebras (gateway) in 2026: Wafer-Scale Speed, Three Models

Cerebras (gateway) drives Atlas on wafer-scale engines: GPT-OSS 120B at $0.35 / $0.75 per Mtok, 131K tokens (131,072) context, and a catalog of three models.

Models

Atlas with Gemini 3.1 Flash Lite: The $0.25 Small Model Slot in 2026

Gemini 3.1 Flash Lite in Atlas: 1,048,576 tokens of context at $0.25 / $1.50 per Mtok, wired to the small_model slot in atlas.json. Setup, limits, and tradeoffs.

Models

Atlas with DeepSeek Reasoner: Chain-of-Thought Debugging at $0.28 in 2026

DeepSeek Reasoner in Atlas: visible reasoning traces on a 1M window at $0.14 / $0.28 per Mtok, roughly 640x below GPT-5.5 Pro's $180 output rate. Setup and limits.

Models

Atlas with Claude Fable 5: Plan Mode Model Guide for 2026

Claude Fable 5 is Anthropic's premium June 2026 model at $10 / $50 per Mtok with a 1M window. Use it for one expensive Atlas planning pass, then drop back down.

Models

Atlas with GPT-5.1 Codex Max: The Cheapest Codex Tier in 2026

GPT-5.1 Codex Max still bills at the November 2025 rate of $1.25 / $10 per Mtok on a 400K window. The lowest-cost entry into OpenAI's Codex post-training for Atlas.

Models

Atlas with GPT-5.1: The Cheapest Full Size GPT-5 Input Price in 2026

GPT-5.1 in Atlas: 400K context, $1.25 per Mtok input and $10 per Mtok output, cheaper on input than every GPT-5 release that followed it in 2026.

Models

Atlas with Qwen2.5 14B Instruct in 2026: The Open-Weights Main Model Tier

Qwen2.5 14B Instruct drives Atlas at $0.35 per Mtok input and $1.40 per Mtok output, fits a single 24 GB GPU at 4-bit, and holds a real multi-file edit plan.

Models

Atlas with GLM-4.5-Air: The 106B Self-Hostable Cheap Slot in 2026

GLM-4.5-Air drives Atlas at $0.20 per Mtok input and $1.10 per Mtok output on a 128K tokens (131,072) window. A 106B total / 12B active MIT-licensed MoE.

Models

Atlas with DeepSeek R1 (0528): The Open Reasoning Trace, 2026

Run Atlas on DeepSeek R1 (0528) in 2026. DeepInfra hosts the MIT-licensed open reasoning model at $0.50 per Mtok in, $2.15 per Mtok out, 160K tokens context.

Models

Atlas with Qwen3 235B-A22B: Flagship Sparse Reasoning in 2026

Qwen3 235B-A22B in Atlas: 235B total parameters, 22B active per token, $0.70 per Mtok input and $2.80 per Mtok output, 128K tokens (131,072) of context.

Models

Atlas with Qwen2.5 72B (local via Ollama): Air-Gapped Coding in 2026

Run Atlas fully offline on Qwen2.5 72B (local via Ollama) in 2026. Free (self-hosted), about 47 GB at Q4_K_M, and a 32,768 token local context cap.

Models

Atlas with Gemini 2.5 Flash: The Workhorse small_model Pick for 2026

Gemini 2.5 Flash in Atlas: reasoning enabled, a 1,048,576 token context, and $0.3 per Mtok input, a sixth of what a frontier tier model charges to read the same repo.

Models

Atlas with GPT-OSS 120B (hosted): picking the right host in 2026

Run Atlas on GPT-OSS 120B (hosted) in 2026. Identical Apache weights cost $0.037 per Mtok on DeepInfra and $0.35 on Cerebras, a 9.5x input spread. Full host guide.

Models

Atlas with GPT-5.4 nano: The $0.20 Background Model in 2026

GPT-5.4 nano is OpenAI's cheapest reasoning-capable model at $0.20 / $1.25 per Mtok with a 400K window. Built for the titles, summaries, and classification Atlas runs constantly.

Models

Atlas with Qwen Turbo: The $0.05 per Mtok Small Model Slot in 2026

Run Atlas on Qwen Turbo in 2026. Alibaba's cheapest reasoning tier gives 1M tokens (1,000,000) of context at $0.05 per Mtok input, $0.20 per Mtok output.

Models

Atlas with Qwen3.5 27B: The Dense Entry Point to Qwen3.5 in 2026

Qwen3.5 27B gives Atlas 256K tokens (262,144) of context and predictable dense latency at $0.30 per Mtok input and $2.40 per Mtok output. Setup, costs, and honest tradeoffs.

Models

Atlas with StarCoder2 (local via Ollama): Auditable Training Data in 2026

StarCoder2 is BigCode's open code model, trained on The Stack v2 with full data provenance. Free self-hosted, 600-plus languages, 16,384 token context. Atlas setup.

Models

Atlas with Perplexity Sonar in 2026: The Live-Web Research Model, Not the Build Model

Atlas with Perplexity Sonar in 2026: live web results with citations, Sonar at $1.00/$1.00 per Mtok, Sonar Pro at 200,000 tokens, and why it is not a build model.

Models

Atlas with Gemma 4 12B (Ollama): 256K Context from a 7.6GB Download in 2026

Gemma 4 12B (Ollama) is a 7.6GB download with a 256K tokens (262,144) context, Free (self-hosted). The longest Gemma window that fits a mid-range GPU. Atlas setup.

Models

Atlas with Poolside Laguna M.1 in 2026

Poolside Laguna M.1 in Atlas, 2026: a coding-native reasoning model with 262,144 tokens of context, free on Poolside's API and $0.20/$0.40 per Mtok on OpenRouter.

Models

Atlas with Devstral Small 2505: The Original Agent-First Model in 2026

Devstral Small 2505 started Mistral's Devstral line in May 2025: 128,000 tokens at $0.10 / 1M input tokens. Atlas setup, why it was superseded, and when to pin it.

Models

Atlas with OpenAI o4-mini (2026): Cheap Reasoning for Parallel Subagents

OpenAI o4-mini drives Atlas at $1.10 per Mtok input and $4.40 per Mtok output on a 200K context. A strict upgrade over o3-mini at identical price, with real tradeoffs.

Models

Atlas with Qwen Plus: 1M Context for $0.40 per Mtok in 2026

Run Atlas on Qwen Plus in 2026. Alibaba's mid tier gives 1M tokens (1,000,000) of context with reasoning at $0.40 per Mtok input, $1.20 per Mtok output.

Models

Atlas with Qwen3.5 397B-A17B: The Qwen3.5 Flagship in 2026

Qwen3.5 397B-A17B is Alibaba's Qwen3.5 flagship: 397B total, 17B active, 256K tokens (262,144) of context, $0.60 per Mtok input and $3.60 per Mtok output, running in Atlas.

Models

Atlas with Command R7B in 2026: Cohere's Cheapest Model, Used Correctly

Command R7B costs $0.0375 per Mtok input, about 1/66th of Command A, and still carries a 128,000 token context. Here is how to slot it into Atlas without wrecking your code.

Models

Atlas with Mixtral 8x22B: The Largest Open MoE of Its Era in 2026

Mixtral 8x22B scaled the MoE idea in April 2024: 64,000 tokens at $2.00 / 1M input tokens and $6.00 / 1M output tokens. Atlas setup, self-hosting, and honest limits.

Models

Atlas with Claude Opus 4.1: The Pre Price Cut Opus in 2026

Claude Opus 4.1 runs Atlas on a 200K context at $15 per Mtok input, $75 per Mtok output, with a 32K output ceiling. Setup, cost warnings, and better alternatives.

Models

Atlas with Gemini 2.0 Flash: A Fast Reader With an 8,192 Token Output Cap in 2026

Gemini 2.0 Flash in Atlas: the December 2024 release with a 1,048,576 token context, $0.1 per Mtok input, no reasoning mode, and an 8,192 token output ceiling.

Models

Atlas with Grok Build 0.1: xAI's First Coding-Agent Model (2026)

Grok Build 0.1 runs in Atlas at $1 / $2 per Mtok with a 256K context and 256,000 max output tokens. A coding-specialized model, but it is a 0.1 release.

Models

Atlas with GPT-5.4 mini: A Reasoning Mini Tier for 2026

GPT-5.4 mini runs $0.75 / $4.50 per Mtok on a 400K window, undercutting Claude Haiku 4.5's $1 with double the context. A reasoning-capable small_model for Atlas.

Models

Atlas with Gemma 4 26B A4B: The Open-Weights MoE Option in 2026

Gemma 4 26B A4B in Atlas: a sparse MoE with about 4B active of 26B total parameters, a 262,144 token context, a 32,768 token output cap, and no public price.

Models

Atlas with Gemini 3.1 Pro Custom Tools: Cost, Context, and Setup in 2026

Run Atlas on Gemini 3.1 Pro Custom Tools in 2026: a 1,048,576 token context, $2 per Mtok input, and a tool-calling checkpoint built for a dense tool surface.

Models

Atlas with MiniMax-M2.5 in 2026: Reasoning Under a Dollar

MiniMax-M2.5 drives Atlas on a 230B efficient-MoE at $0.30 per Mtok input and $1.20 per Mtok output, with a 204,800 token context and 131,072 max output tokens.

Models

Atlas with SiliconFlow in 2026: Qwen3 Coder 480B at $0.25/$1.00 per Mtok

Atlas with SiliconFlow in 2026: Qwen3-Coder-480B-A35B at $0.25/$1.00 per Mtok, a free Qwen3.5-4B small_model, and the api.siliconflow.cn residency question.

Models

Atlas with Llama 4 Scout (Ollama): A 10M-Token Window on Local Hardware in 2026

Run Atlas on Llama 4 Scout (Ollama): a 67GB 16-expert MoE with a 10M-token context and image input, free self-hosted. The memory math before you pull 67GB.

Models

Atlas with DeepCoder 14B (Ollama): RL-Tuned for First-Attempt Diffs in 2026

DeepCoder 14B (Ollama) is a 9.0GB RL-tuned coder with a 128K context, built for first-attempt correctness. Free (self-hosted). Atlas setup and tradeoffs for 2026.

Models

Atlas with NVIDIA Nemotron 3 Ultra 550B A55B in 2026

Run Atlas on NVIDIA Nemotron 3 Ultra 550B A55B in 2026. A 550B mixture of experts with 55B active, 1,000,000 tokens on NVIDIA NIM at $0.50/$2.50 per Mtok.

Models

Atlas with Phi-4 Mini 3.8B (Ollama): Native Function Calling at 2.5GB in 2026

Phi-4 Mini 3.8B (Ollama) is a 2.5GB model with 128K tokens (131,072) of context and native function calling, Free (self-hosted). The Atlas small_model that can route tools.

Models

Atlas with MiniMax-M2: The $0.30 Open-Weights Agent Model in 2026

MiniMax-M2 runs Atlas at $0.30 per Mtok input and $1.20 per Mtok output with a 196,608 token context, open weights on HuggingFace, and an Anthropic-compatible API.

Models

Atlas with Mistral Small 24B (Ollama): The Single-GPU Commercial Pick for 2026

Run Atlas on Mistral Small 24B (Ollama): 14GB weights, 32K context, Apache 2.0, free self-hosted. Setup, the 32K vs 128K tag trap, and when a coder beats it.

Models

Atlas with Liquid AI LFM2-24B-A2B in 2026

Liquid AI LFM2-24B-A2B in Atlas, 2026: a liquid neural network MoE at $0.03/$0.12 per Mtok on Together AI, with a 32,768 token context and matching output.

Models

Atlas with GitHub Models in 2026: Free Model Access Behind a GITHUB_TOKEN

Run Atlas on GitHub Models in 2026: every model listed at $0/$0 per Mtok, auth with the GITHUB_TOKEN you already have, and 256,000 tokens on AI21 Jamba 1.5 Large.

Models

Atlas with Baseten in 2026: 262,000 Tokens In, 262,000 Tokens Out

Atlas with Baseten in 2026: Kimi K2.7 Code at 262,000 tokens in and 262,000 out, GPT OSS 120B at $0.10/$0.50 per Mtok, and the tradeoffs of dedicated hosting.

Models

Atlas with Llama 3.2 3B (Ollama): The CPU-Only Floor for a Local Agent in 2026

Run Atlas on Llama 3.2 3B (Ollama): a 2.0GB model with a 128K context, free self-hosted. The realistic floor for an Atlas setup with no GPU at all.

Models

Atlas with IBM Granite 4.0 H Micro in 2026

IBM Granite 4.0 H Micro in Atlas, 2026: a hybrid Mamba-Transformer model at $0.017/$0.112 per Mtok on Cloudflare Workers AI, the cheapest input in the registry.

Models

Atlas with Mistral Nemo 12B (local via Ollama): The 12GB GPU Pick for 2026

Mistral Nemo 12B (local via Ollama) in Atlas for 2026: a 7.1GB pull, Free (self-hosted), Tekken tokenizer, and why the KV cache, not the weights, caps context.

Models

Atlas with Mistral Large 3 (2512): Big Diffs, EU Hosted, 2026

Mistral Large 3 (2512) drives Atlas with a 262,144 token context and a matching 262,144 token output at $0.50 / 1M input tokens and $1.50 / 1M output tokens.

Models

Atlas with CodeGemma 7B (Ollama): Fill-in-the-Middle on 8GB in 2026

CodeGemma 7B (Ollama) is Google's 5.0GB code model with fill-in-the-middle training and an 8K context, Free (self-hosted). Atlas setup and honest tradeoffs for 2026.

Models

Atlas with IBM Granite 3.3 8B (Ollama): the Free Local small_model for 2026

IBM Granite 3.3 8B (Ollama) is a 4.9GB general model with 128K tokens (131,072) of context, Free (self-hosted). Assign it to small_model in Atlas. Setup and limits.

Models

Atlas with GLM-5-Turbo: Setup, Pricing, and Tradeoffs in 2026

Run Atlas on GLM-5-Turbo from Z.ai. A 200,000 token context at $1.20 per Mtok input and $4.00 per Mtok output, plus why Turbo costs more than base GLM-5.

Models

Atlas with Mistral Large 2.1 (2411): A 2026 Setup Guide

Mistral Large 2.1 (2411) runs Atlas on EU infrastructure with a 131,072 token context at $2.00 / 1M input tokens and $6.00 / 1M output tokens. Setup and honest limits.

Models

Atlas with Gemini Flash Latest: The Rolling Alias Explained (2026)

gemini-flash-latest in Atlas: a rolling alias, not a pinned checkpoint. $0.3 per Mtok input, $2.5 per Mtok output, 1,048,576 token context, and no reproducibility.

Models

Atlas with Kimi K2 0711: The Original Trillion-Parameter Preview in 2026

Run Atlas on Kimi K2 0711 in 2026. Moonshot's original K2 preview costs $0.60 per Mtok input, $2.50 per Mtok output, with a 128K tokens (131,072) context.

Models

Atlas with Qwen3 32B (Ollama): dense reasoning over raw speed in 2026

Qwen3 32B (Ollama) in Atlas: dense 20GB weights on a 24GB card, 40K tokens (40,960) of context, Free (self-hosted). Slower than the MoE, steadier on hard problems.

Models

Atlas with Qwen3-Next 80B-A3B Thinking: The Reasoning Tier in 2026

Drive Atlas with Qwen3-Next 80B-A3B Thinking: a reasoning trace over 128K tokens (131,072) of context at $0.50 per Mtok input and $6.00 per Mtok output.

Models

Atlas with Qwen3.6 Plus: The Stable Million-Token Tier in 2026

Qwen3.6 Plus gives Atlas 1M tokens (1,000,000) of context at $0.50 per Mtok input and $3.00 per Mtok output, matching Qwen3.7 Plus while keeping 3.6 generation behavior.

Models

Atlas with Kimi K2 Turbo: Pricing, Context, and Setup in 2026

Kimi K2 Turbo in Atlas costs $2.40 per Mtok input and $10.00 per Mtok output on a 256K tokens (262,144) window. A latency purchase, not a capability upgrade.

Models

Atlas with DeepSeek V3 (open weights): The Frozen 671B Baseline in 2026

Run Atlas on DeepSeek V3 (open weights) in 2026. DeepInfra hosts the MIT-licensed 671B MoE at $0.32 per Mtok input, $0.89 per Mtok output, 128K context.

Models

Atlas with DeepSeek V3.1 (open weights): Togglable Thinking in 2026

Run Atlas on DeepSeek V3.1 (open weights) in 2026. One MIT-licensed checkpoint with thinking and non-thinking modes, $0.25 per Mtok in and $0.95 per Mtok out.

Models

Atlas with GLM-5.2: A 1M Token Open-Weights Model at $1.40 per Mtok (2026)

GLM-5.2 drives Atlas on a 1M token context at $1.40 / $4.40 per Mtok. Z.ai's June 2026 flagship, the first GLM to reach 1M, with open-weights lineage.

Models

Atlas with Qwen3-Coder Next (local via Ollama): the Top-Ranked Local Coder in 2026

Qwen3-Coder Next is the top-ranked local coding model of mid-2026: 262,144 token context, free self-hosted, or $0.22 / $1.80 per Mtok on Bedrock. Atlas setup and tradeoffs.

Models

Atlas with Upstage Solar in 2026

Run Atlas on the Upstage Solar lineup in 2026: solar-pro3 and solar-pro2 at $0.25/$0.25 per Mtok, solar-mini at $0.15/$0.15, across three context sizes.

Models

Atlas with Qwen3 30B-A3B (Ollama): the MoE throughput trade in 2026

Qwen3 30B-A3B (Ollama) in Atlas: 30B total parameters, roughly 3B active per token, 19GB of weights, 256K tokens (262,144) of context, Free (self-hosted).

Models

Atlas with Llama 3.1 8B (Ollama): 128K Context on an 8GB Card in 2026

Run Atlas on Llama 3.1 8B (Ollama): Meta's 4.9GB workhorse with a 128K context, free self-hosted. Setup, the KV cache catch, and when a coder model wins.

Models

Atlas with Gemma 2 27B (Ollama): the Free Local Diff Reviewer in 2026

Gemma 2 27B (Ollama) is Google's 2024 flagship open model, 16GB and Free (self-hosted), with 8K tokens (8,192) of context. Use it to review Atlas diffs, not write them.

Models

Atlas with Phi-3 Medium 14B (Ollama): 128K Context at 7.9GB in 2026

Phi-3 Medium 14B (Ollama) gives Atlas a 128K tokens (131,072) window from a 7.9GB download, Free (self-hosted). Pull phi3:14b, not :medium-4k. Setup and tradeoffs.

Models

Atlas with Gemma 4 31B (Ollama): the Flagship Gemma 4 Tag in 2026

Gemma 4 31B (Ollama) is the largest Gemma 4 tag: 20GB weights, a 256K tokens (262,144) context, Free (self-hosted), plus a 31b-coding-mtp-bf16 variant. Atlas setup.

Models

Atlas with Qwen2.5-Coder 7B (Ollama): the default fully offline setup for 2026

Qwen2.5-Coder 7B (Ollama) in Atlas: the 4.7GB default tag with 32K tokens (32,768) of context, Free (self-hosted), and no cloud key anywhere in the agent loop.

Models

Atlas with Gemma 3 27B Instruct on Amazon Bedrock: Flat Pricing, 202,752 Tokens (2026)

Gemma 3 27B Instruct in Atlas via Amazon Bedrock: $0.12 per Mtok input, $0.2 per Mtok output, a 202,752 token context, open weights, and an 8,192 token output cap.

Models

Atlas with Gemma 3 4B Instruct: A Triage Model, Not a Builder (2026)

Gemma 3 4B Instruct in Atlas via Amazon Bedrock: $0.04 per Mtok input, $0.08 per Mtok output, a 128K context, and a 4,096 token output cap that rules out diffs.

Models

Atlas with GPT-5: The Original 400K Reasoning Model in 2026

GPT-5 in Atlas: the August 2025 launch model with a 400K context, 128K max output, and $1.25 per Mtok input, still the floor price for a full size GPT-5 class model.

Models

Atlas with Command A: Cohere's 256K Context Flagship in 2026

Command A gives Atlas a 256,000 token read window at $2.5 per Mtok input and $10 per Mtok output, with an 8,000 token output cap that shapes how you refactor.

Models

Atlas with North Mini Code in 2026: A 64,000 Token Output Budget

North Mini Code gives Atlas a 256,000 token context and a 64,000 token output, 8x Command A's 8,000, listed at $0 per Mtok on both input and output in the registry.

Models

Atlas with Command A Reasoning: Reasoning You Can Deploy On-Prem (2026)

Command A Reasoning gives Atlas a 256K window at $2.50 / $10 per Mtok, and Cohere lets you run it on-prem or in a VPC. Reasoning with no price premium.

Models

Atlas with Groq (gateway) in 2026: LPU Speed for the Agent Loop

Groq (gateway) runs open models on LPU hardware for Atlas: GPT-OSS 120B at $0.15 / $0.60 per Mtok, 131K tokens (131,072) context, and no Claude or GPT-5.

Models

Atlas with Qwen2.5 7B Instruct in 2026: The Cheap Dense Small Model

Qwen2.5 7B Instruct runs Atlas's small_model slot at $0.175 per Mtok input and $0.70 per Mtok output, keeping the full 131,072 token window on a dense 7B checkpoint.

Models

Atlas with Kimi K2.6: The Generalist Reasoning Tier in 2026

Kimi K2.6 drives Atlas at $0.95 per Mtok input and $4.00 per Mtok output on a 256K tokens (262,144) window. The generalist pick when the job is not purely code.

Models

Atlas with Kimi K2.7 Code HighSpeed: The 2x Throughput Tier in 2026

Kimi K2.7 Code HighSpeed runs Atlas at $1.90 per Mtok input and $8.00 per Mtok output, exactly 2x K2.7 Code, on the same 256K tokens (262,144) context window.

Models

Atlas with Qwen3-Coder 480B-A35B Instruct: The Open Frontier Coder in 2026

Qwen3-Coder 480B-A35B Instruct in Atlas: 480B total parameters, 35B active per token, 262,144 tokens of context, $1.50 per Mtok in and $7.50 per Mtok out.

Models

Atlas with Mistral Small 4 (2603): Cheap Reasoning in 2026

Mistral Small 4 (2603) brings reasoning to the Small tier: 256,000 tokens at $0.15 / 1M input tokens and $0.60 / 1M output tokens. Atlas setup, costs, tradeoffs.

Models

Atlas with DeepSeek V3.2 (open weights): Sparse Attention at $0.38 Output, 2026

Run Atlas on DeepSeek V3.2 (open weights) in 2026. DeepSeek Sparse Attention gives 160K tokens (DeepInfra) at $0.26 per Mtok in and $0.38 per Mtok out.

Models

Atlas with Grok 4.20 (Reasoning) in 2026: A 1M Token Reader

Grok 4.20 (Reasoning) reads 1,000,000 tokens at $1.25 per Mtok input and writes at $2.5 per Mtok, but caps output at 30,000 tokens. A superb reader, a terse writer.

Models

Atlas with Poolside Laguna XS 2.1 in 2026

Poolside Laguna XS 2.1 in Atlas, 2026: the fast tier of Poolside's coding-native line at $0.06/$0.12 per Mtok on OpenRouter, holding 262,144 tokens of context.

Models

Atlas with NVIDIA Nemotron 3 Nano 30B A3B in 2026

Nemotron 3 Nano 30B A3B in Atlas, 2026: 3B active parameters at $0.05/$0.20 per Mtok on DeepInfra, free on NVIDIA NIM, up to 1,048,576 tokens on Ollama Cloud.

Models

Atlas with Command R 35B (Ollama): A RAG-Native Model for Retrieval-Heavy Work in 2026

Run Atlas on Command R 35B (Ollama): Cohere's 19GB RAG and tool-use model with a 128K context, free self-hosted. Check the research license before commercial use.

Models

Atlas with IBM Granite Code 20B (Ollama): More Capacity, Less Window in 2026

IBM Granite Code 20B (Ollama) is 12GB on disk and Free (self-hosted), but the 20b tag drops to 8K tokens (8,192) where the 8B instruct advertises 125K.

Models

Atlas with MiniMax-M2.7-highspeed: The Fast Lane in 2026

MiniMax-M2.7-highspeed gives Atlas priority serving on MiniMax's newest 230B MoE at $0.60 per Mtok input and $2.40 per Mtok output, with a 204,800 token context.

Models

Atlas with MiniMax-M3 in 2026: A Million Tokens for Thirty Cents

MiniMax-M3 gives Atlas a 1,000,000 token context at $0.30 per Mtok input and $1.20 per Mtok output, the cheapest large-context reasoning offer of 2026. Setup and limits.

Models

Atlas with Mixtral 8x7B (Ollama): Sparse MoE Throughput in 2026

Run Atlas on Mixtral 8x7B (Ollama): a 26GB Apache 2.0 sparse mixture of experts with 32K context, free self-hosted. Memory math, throughput, and honest limits.

Models

Atlas with Llama 3.1 8B (local via Ollama): The 4.9GB Baseline for 2026

Llama 3.1 8B (local via Ollama) in Atlas for 2026: a 4.9GB pull that fits 8GB of VRAM, Free (self-hosted), and honest limits on a general-purpose 8B model.

Models

Atlas with Devstral Medium (2507): Agent-Trained Frontier Coding in 2026

Devstral Medium (2507) gives Atlas agent-first training at $0.40 / 1M input tokens and $2.00 / 1M output tokens, with 128,000 tokens in and out. Setup and tradeoffs.

Models

Atlas with Mistral Nemo: The 128K Small Model Slot in 2026

Running Atlas on Mistral Nemo in 2026: 128,000 tokens of context in a 12B model at $0.15 / 1M input tokens, with the Tekken tokenizer that compresses code.

Models

Atlas with Magistral 24B (Ollama): A Local Reasoning Model for the Plan Agent in 2026

Run Atlas on Magistral 24B (Ollama): Mistral's 14GB reasoning model with a 39K context, free self-hosted. Use it as the plan agent, then hand edits to a coder.

Models

Atlas with Qwen3-Next 80B-A3B Instruct: Setup, Cost, and Tradeoffs in 2026

Run Atlas, the terminal-native AI coding agent, on Qwen3-Next 80B-A3B Instruct: 128K tokens (131,072) of context at $0.50 per Mtok input and $2.00 per Mtok output.

Models

Atlas with Command R+ in 2026: Stronger Tool Use, 128K Context

Command R+ drives Atlas with a 128,000 token context and stronger multi step tool use, priced at $2.5 per Mtok input and $10 per Mtok output with a 4,000 token cap.

Models

Atlas with Qwen3.5 Plus: A Million-Token Window for $0.40 per Mtok in 2026

Qwen3.5 Plus gives Atlas a 1M tokens (1,000,000) context window at $0.40 per Mtok input and $2.40 per Mtok output. What the million tokens buy, and what closed weights cost.

Models

Atlas with Qwen3.6 27B: The 2026 Dense Checkpoint, Priced Honestly

Qwen3.6 27B is the dense reasoning model of Alibaba's April 2026 line: 256K tokens (262,144) of context at $0.60 per Mtok input and $3.60 per Mtok output, running inside Atlas.

Models

Atlas with DeepSeek V4 Pro: 384K Output Tokens at $0.87 in 2026

DeepSeek V4 Pro in Atlas: the April 2026 flagship with a 1M token window, 384,000 max output tokens, and $0.435 / $0.87 per Mtok. Setup, data residency, tradeoffs.

Models

Atlas with Amazon Nova Micro in 2026: The Cheapest Model on Bedrock

Amazon Nova Micro costs $0.035 per Mtok input, the lowest price in the Bedrock catalog, with a 128K token context. Use it as Atlas's small_model, never as the build loop.

Models

Atlas with Command A Plus: Run a Frontier Model Inside Your Own VPC (2026)

Command A Plus drives Atlas at $2.50 / $10 per Mtok on a 128K window, and Cohere can deploy it in your own VPC or on-prem so code never leaves your network.

Models

Atlas with Gemma 4 31B IT: Running Google's Open Weights Locally in 2026

Gemma 4 31B IT in Atlas: Google's April 2026 open-weights release with a 262,144 token window, free self-hosted via ollama pull gemma4:31b, or $0.99 / $1.49 on Cerebras.

Models

Atlas with GPT-OSS 120B (local via Ollama): Offline Reasoning for Air-Gapped Work in 2026

GPT-OSS 120B is the strongest fully offline reasoning model for Atlas: a 65GB MXFP4 download, 131,072 token context, free self-hosted, or $0.15 / $0.60 per Mtok on Groq.

Models

Atlas with Cloudflare Workers AI in 2026: The Cheapest Input Token in the Registry

Atlas with Cloudflare Workers AI in 2026: IBM Granite 4.0 H Micro at $0.017/$0.112 per Mtok, Kimi K2.7 Code at 262,144 tokens, and edge inference tradeoffs.

Models

Atlas with NVIDIA NIM: Free-Tier Open Weights and the Nemotron Home Turf in 2026

Run Atlas on NVIDIA NIM: most endpoints listed at $0/$0 per Mtok, Nemotron 3 Ultra 550B at 1,000,000 tokens for $0.50/$2.50. Setup, rate limits, and catalog noise.

Models

Atlas with Devstral Small 2 24B (local via Ollama): the Agent-First Local Model in 2026

Devstral Small 2 24B is Mistral's agent-first local model: a 14GB Ollama download, 128K context, free self-hosted, and it runs on a 16GB GPU. Atlas setup and tradeoffs.

Models

Atlas with Ministral 8B: The Cheap Slot That Can Still Call Tools in 2026

Ministral 8B runs 128,000 tokens at $0.10 / 1M input tokens and $0.10 / 1M output tokens, symmetric. The Atlas small_model upgrade when subagents misfire on schemas.

Models

Atlas with DeepSeek-R1 14B Distill (Ollama): The Local Plan Agent in 2026

DeepSeek-R1 14B Distill (Ollama) is 9.0GB, roughly 11GB to serve, with a 128K context. Drive the Atlas plan agent on a 12GB card in 2026. Free (self-hosted).

Models

Atlas with Mistral NeMo 12B (Ollama): 128K Context on a 12GB Card in 2026

Run Atlas on Mistral NeMo 12B (Ollama): 7.1GB, a 128K practical context, free self-hosted. Why the Ollama tag says 1000K, and what limit.context to actually set.

Models

Atlas with MiniMax-M2.7 in 2026: Agentic Reasoning at $0.30

MiniMax-M2.7 is MiniMax's March 2026 agentic 230B MoE. It runs Atlas at $0.30 per Mtok input and $1.20 per Mtok output with a 204,800 token context and 131,072 output.

Models

Atlas with Qwen2.5-Coder 14B (Ollama): a real local build agent in 2026

Qwen2.5-Coder 14B (Ollama) in Atlas: 9.0GB of Q4_K_M weights, roughly 11GB to serve, 32K tokens (32,768) of context, Free (self-hosted), steady on tool chains.

Models

Atlas with Claude Sonnet 4.5: The First 1M Token Claude in 2026

Claude Sonnet 4.5 gives Atlas a 1M token window at $3 per Mtok input, $15 per Mtok output. Setup, the 64K output ceiling, and when Sonnet 5 is the better pin.

Models

Atlas with GLM-4.7 Flash: A Free 200K Context Model for the small_model Slot (2026)

GLM-4.7 Flash is free at $0 / $0 per Mtok with a 200,000 token context. Set it as Atlas's small_model so titles and summaries cost nothing at all.

Models

Atlas with Llama 4 Maverick: 1M Context on Open Weights in 2026

Llama 4 Maverick gives Atlas a 1M token context on open weights at $0.24 / $0.97 per Mtok on Bedrock or $0.20 / $0.80 on DeepInfra. Setup, tradeoffs, and fit.

Models

Atlas with Claude Haiku 4.5: The Cheap Slot in 2026

Claude Haiku 4.5 runs Atlas's small_model slot at $1 / $5 per Mtok with a 200K window. Titles, commit summaries, and cheap subagent fan-out, priced honestly for 2026.

Models

Atlas with Qwen3.7 Max: Alibaba's May 2026 Frontier Tier at $2.50 / $7.50

Qwen3.7 Max in Atlas: Alibaba's May 2026 flagship, a 1M context model at $2.50 / $7.50 per Mtok, or $1.25 / $3.75 through Together. Setup with DASHSCOPE_API_KEY.

Models

Atlas with Devstral 2: Mistral's Agent-First Coding Model in 2026

Devstral 2 in Atlas: Mistral's agent-trained coding model with a 262,144 token window at $0.40 / $2 per Mtok, loaded through @ai-sdk/mistral. Setup and tradeoffs.

Models

Atlas with GLM-4.5-Flash: The Free Reasoning Slot in 2026

GLM-4.5-Flash is listed at $0.00 per Mtok input and output on Z.ai, with the full 128K tokens (131,072) context. Free background traffic for Atlas, rate limited.

Models

Atlas with Gemini 2.0 Flash-Lite: The Cheapest Google Model in the Registry (2026)

Gemini 2.0 Flash-Lite in Atlas: $0.075 per Mtok input, $0.3 per Mtok output, a 1,048,576 token context, an 8,192 token output cap, and no reasoning mode.

Models

Atlas with OpenAI o1-pro (2026): The $600 Per Mtok Question

OpenAI o1-pro is the most expensive model in the OpenAI registry at $150 per Mtok input and $600 per Mtok output. Here is what it does in Atlas and why o3 usually wins.

Models

Atlas with NVIDIA Nemotron 3 Super 120B A12B in 2026

Nemotron 3 Super 120B A12B in Atlas, 2026: a reasoning MoE from $0.15/$0.65 per Mtok on Vercel AI Gateway, with 262,144 tokens of context on NVIDIA NIM.

Models

Atlas with Nebius Token Factory in 2026: EU Infrastructure and the -fast Latency Lever

Atlas with Nebius Token Factory in 2026: EU-operated infrastructure, Qwen3.5-397B-A17B at $0.60/$3.60 per Mtok, and an 8,192 token output cap to plan around.

Models

Atlas with Gemma 4 E4B (Ollama): the Default Gemma 4 Tag in 2026

Gemma 4 E4B (Ollama) is the :latest tag of Google's newest Gemma line: 9.6GB, 128K tokens (131,072), Free (self-hosted), with qat, mlx, mxfp8 and nvfp4 quants.

Models

Atlas with Llama 3.1 70B (local via Ollama): A 43GB Private Agent in 2026

Llama 3.1 70B (local via Ollama) in Atlas for 2026: a 43GB pull for a 48GB GPU or 64GB Mac, Free (self-hosted), with 128,000 tokens of context on your own hardware.

Models

Atlas with Mixtral 8x7B: Running the Original Sparse MoE in 2026

Mixtral 8x7B put sparse mixture-of-experts on the map in December 2023. In Atlas it gives 32,000 tokens at $0.70 / 1M input tokens. Setup, limits, and honest fit.

Models

Atlas with Qwen3 Coder Flash: 1M Context at $0.30 per Mtok in 2026

Qwen3 Coder Flash in Atlas: 1M tokens (1,000,000) of context, 65,536 token output, $0.30 per Mtok input and $1.50 per Mtok output, tuned for fast edit loops.

Models

Atlas with Qwen3 8B (local via Ollama): The Laptop Offline Setup for 2026

Qwen3 8B (local via Ollama) is the easiest offline Atlas setup in 2026: roughly 5.2 GB at Q4_K_M, Free (self-hosted), and a thinking model on a laptop GPU.

Models

Atlas with GPT-4.1 nano (2026): A $0.10 Triage Model, Not a Coding Agent

GPT-4.1 nano runs in Atlas at $0.10 per Mtok input and $0.40 per Mtok output on a 1,047,576 token window. Use it for file triage and commit messages, not refactors.

Models

Atlas with GPT-OSS 20B (local via Ollama): Local Reasoning on a 16GB Card in 2026

GPT-OSS 20B is OpenAI's open-weight reasoning model: 131,072 token context, runs on a single 16GB GPU, free self-hosted, or $0.075 / $0.30 per Mtok via Groq.

Models

Atlas with DeepSeek-R1 (local via Ollama): Visible Reasoning at Every Size in 2026

DeepSeek-R1 runs locally from 1.5B to 671B under one Ollama tag, with a 128K to 164K context, free self-hosted, or $1.35 / $5.40 per Mtok via Bedrock. Atlas setup.

Models

Atlas with GLM-4.6: 200K Open Weights at the 4.5 Price in 2026

GLM-4.6 runs Atlas on a 200K tokens (204,800) context at $0.60 per Mtok input and $2.20 per Mtok output, the same price GLM-4.5 charged on a 128K window.

Models

Atlas with Qwen3.5 122B-A10B: The Middle MoE, Reviewed for 2026

Qwen3.5 122B-A10B brings 10B active parameters and 256K tokens (262,144) of context to Atlas at $0.40 per Mtok input and $3.20 per Mtok output. Setup, value, and where it loses.

Browse all Atlas resource hubs