Resource hub
Models and providers
Pick and configure the model behind Atlas: frontier models, fast models, open weights, and fully local inference.
Atlas with GPT-5.1 Codex: Plan First, Then Build in 2026
GPT-5.1 Codex in Atlas: a frontier model from OpenAI at $1.25 per Mtok input, $10 per Mtok output, built to execute an approved plan across a long tool chain.
Atlas with Qwen3 235B-A22B (local via Ollama): Self-Hosted Flagship in 2026
Qwen3 235B-A22B (local via Ollama) in Atlas: roughly 140 GB at 4-bit, Free (self-hosted), 22B active of 235B total, and a brutal memory floor to plan for.
Atlas with GPT-5.3 Codex Spark: Fast Interactive Edits in 2026
GPT-5.3 Codex Spark is the low latency Codex model: 128K context, 32K max output, $1.75 per Mtok input, $14 per Mtok output. Atlas setup and honest tradeoffs.
Atlas with Qwen3.5 35B-A3B: The Cheapest Reasoning Tier of 2026
Qwen3.5 35B-A3B is the cheapest reasoning tier in the Qwen3.5 line at $0.25 per Mtok input and $2.00 per Mtok output, with 256K tokens (262,144) of context for Atlas.
Atlas with Qwen2.5 32B Instruct: Cost, Context, and Setup in 2026
Run Atlas on Qwen2.5 32B Instruct in 2026. 128K tokens (131,072) of context, $0.70 per Mtok input, $2.80 per Mtok output, and an offline Ollama route.
Atlas with GLM-4.7: The Value Pick of the GLM Line in 2026
GLM-4.7 drives Atlas at $0.60 per Mtok input and $2.20 per Mtok output on a 200K tokens (204,800) window, undercutting Kimi K2.6's $4.00 output rate.
Atlas vs Mistral Vibe for Code: Terminal AI Coding Agents in 2026
Compare Atlas and Mistral Vibe for Code in 2026. Atlas offers terminal-native TUI, explicit diffs, and BYO models. Mistral Vibe for Code provides a four-model stack, multi-platform access, and EU data sovereignty.
Atlas with DeepSeek Coder 6.7B (Ollama): the thin-hardware fallback in 2026
DeepSeek Coder 6.7B (Ollama) in Atlas: a 3.8GB pull, roughly 6GB to serve, 16K tokens (16,384) of context, Free (self-hosted). Dated, but it starts fast.
Atlas with Llama 3.2 1B (local via Ollama): the 2026 offline small_model slot
Llama 3.2 1B (local via Ollama) in Atlas: a 1.3GB CPU-friendly pull with a 128,000 token window, Free (self-hosted), best wired to small_model, not model.
Atlas with GPT-5.3 Chat: The small_model Slot Pick for 2026
GPT-5.3 Chat is a non reasoning 128K model at $1.75 per Mtok input, $14 per Mtok output. Why it belongs in the Atlas small_model slot, not the main build loop.
Atlas with Qwen Max: Setup, Cost, and the 32,768 Token Limit in 2026
Run Atlas on Qwen Max in 2026. Alibaba's original flagship costs $1.60 per Mtok input, $6.40 per Mtok output, with a 32K tokens (32,768) context window.
Atlas with GLM-4.5: The MIT-Licensed Open Weights Baseline in 2026
GLM-4.5 runs Atlas at $0.60 per Mtok input and $2.20 per Mtok output on a 128K tokens (131,072) window, under an MIT license that permits commercial self-hosting.
Atlas with GPT-5.2 Pro: The Deep Reasoning Tier in 2026
Running Atlas on GPT-5.2 Pro in 2026: a 400K token context, $21 per Mtok input, $168 per Mtok output, and the one shot questions that justify the price.
Atlas with AllenAI Olmo 3 32B Think: the fully open reasoning model in 2026
Run Atlas on AllenAI Olmo 3 32B Think in 2026: 65,536 token context, OpenRouter $0.15/$0.50 per Mtok, and the only weights, data, and training code you can audit.
Atlas with GLM-5: Z.ai's Frontier Tier for Coding Agents in 2026
GLM-5 runs Atlas at $1.00 per Mtok input and $3.20 per Mtok output on a 200K tokens (204,800) window. Z.ai's February 2026 jump past the 4.x line.
Atlas with Gemini 2.5 Flash-Lite: A 1M Context Helper Model for $0.1 per Mtok in 2026
Gemini 2.5 Flash-Lite in Atlas: $0.1 per Mtok input, $0.4 per Mtok output, a 1,048,576 token context, and reasoning enabled. Google's cheapest million token model.
Atlas with MiniMax-M2.5-highspeed in 2026: Paying 2x for Latency
MiniMax-M2.5-highspeed runs Atlas at $0.60 per Mtok input and $2.40 per Mtok output, exactly double base M2.5, for identical weights and a 204,800 token context.
Atlas with Claude Opus 4.5: The Price Break Opus in 2026
Claude Opus 4.5 runs Atlas on a 200K context at $5 per Mtok input, $25 per Mtok output, down from $15/$75 on Opus 4.1. Setup, extended thinking, and tradeoffs.
Atlas with Grok 4.20 (Non-Reasoning) in 2026: Fast, Predictable Edit Passes
Grok 4.20 (Non-Reasoning) bills no thinking tokens, so cost per turn is fully predictable at $1.25 / $2.5 per Mtok, with the same 1,000,000 token context as the reasoning variant.
Atlas with Claude Opus 4.6: Setup, Cost, and Tradeoffs in 2026
Claude Opus 4.6 gives Atlas a 1M token context at $5 per Mtok input, $25 per Mtok output. Setup steps, cost math, and when to pick a newer Opus instead.
Atlas with GPT-5.6 Luna: Best Price Per Context in 2026
GPT-5.6 Luna keeps the full 1,050,000 token window at $1 / $6 per Mtok, one fifth of GPT-5.6. The best price-per-context ratio for agentic coding in Atlas in 2026.
Atlas with Qwen3.6 Max Preview: The Top 3.6 Tier, Reviewed in 2026
Qwen3.6 Max Preview is Alibaba's top Qwen3.6 tier at $1.30 per Mtok input and $7.80 per Mtok output, with 256K tokens (262,144) of context. What it is worth inside Atlas.
Atlas with GPT-5.5 Pro: One Hard Question at a Time in 2026
GPT-5.5 Pro is OpenAI's April 2026 maximum-effort reasoning tier at $30 / $180 per Mtok on a 1,050,000 token window. Invoke it deliberately in Atlas, then switch back.
Atlas with Qwen3 14B: The Middle Dense Tier Worth Pinning in 2026
Qwen3 14B in Atlas for 2026: reasoning-enabled dense 14B at $0.35 per Mtok input and $1.40 per Mtok output, 128K tokens (131,072), roughly 9 GB quantized.
Atlas with Llama 3.1 405B: The 243GB Landmark in 2026
Llama 3.1 405B in Atlas for 2026: 405 billion openly released parameters, a 128,000 token window, Free (self-hosted), and a 243GB download that decides everything.
Atlas with Qwen3.7 Plus: The Newest Million-Token Qwen Tier, June 2026
Qwen3.7 Plus is the value play in Alibaba's 3.7 line: 1M tokens (1,000,000) of context at $0.50 per Mtok input and $3.00 per Mtok output, one fifth the input price of Qwen3.7 Max.
Atlas with Gemini 3 Flash (2026): The $0.5 Default for All Day Agent Loops
Gemini 3 Flash gives Atlas a 1,048,576 token context and 65,536 token output at $0.5 per Mtok input and $3 per Mtok output. The economical default for agent loops.
Atlas with GPT-5.6 Sol: The High Effort GPT-5.6 Variant in 2026
GPT-5.6 Sol tops the 5.6 line at $5 per Mtok input, $30 per Mtok output on a 1,050,000 token window. Atlas setup, small_model routing, and when Sol is overkill.
Atlas with Qwen3 14B (Ollama): the local planning model for 2026
Qwen3 14B (Ollama) in Atlas: dense 9.3GB weights, roughly 11GB to serve, 40K tokens (40,960) of context, Free (self-hosted), better at architecture than raw diffs.
Atlas with Mistral 7B v0.3 (Ollama): The Predictable Local Baseline in 2026
Run Atlas on Mistral 7B v0.3 (Ollama): a 4.4GB Apache 2.0 model with a 32K context, free self-hosted. Setup, tradeoffs, and when to pick a stronger coder.
Atlas with GPT-5.2 Codex: Agentic Coding at $1.75 per Mtok in 2026
GPT-5.2 Codex in Atlas: 400K context, $1.75 per Mtok input and $14 per Mtok output, with Codex post training for long running software engineering work.
Atlas with DeepSeek-R1 32B Distill (Ollama): Local Reasoning on a 24GB Card in 2026
DeepSeek-R1 32B Distill (Ollama) is 20GB of weights with a 128K context, the strongest R1 distill that fits a 24GB consumer GPU. Atlas setup and tradeoffs for 2026.
Atlas with Grok 4.5: A Symmetric 500K Context and Output Budget (2026)
Grok 4.5 drives Atlas at $2 / $6 per Mtok with a 500K token context and a matching 500,000 max output tokens. Atlas routes xAI through the Responses API.
Atlas with GPT-5.6: OpenAI's 2026 Flagship in the Terminal
GPT-5.6 gives Atlas a 1,050,000 token context at $5 / $30 per Mtok. Atlas routes OpenAI through the Responses API so reasoning persists across tool calls in 2026.
Atlas with GPT-4.1 (2026): A Million Token Window at $2 In, $8 Out
GPT-4.1 gives Atlas a 1,047,576 token context at $2 per Mtok input and $8 per Mtok output. Fast, non reasoning, with a 32,768 token output cap. Setup and tradeoffs.
Atlas with Qwen2.5-Coder 32B (Ollama): the air-gapped ceiling in 2026
Qwen2.5-Coder 32B (Ollama) in Atlas: 20GB of weights, roughly 22GB to serve on a 24GB card, 32K tokens (32,768) of context, Free (self-hosted), fully air-gapped.
Atlas with GPT-5.2: The 400K General Purpose GPT in 2026
GPT-5.2 runs Atlas on a 400K context with 128K max output at $1.75 per Mtok input, $14 per Mtok output. Setup, the price rise from GPT-5.1, and the Codex tradeoff.
Atlas with Kimi K2.5: The Cheap Reasoning Default for 2026
Kimi K2.5 gives Atlas a current-generation reasoning model at $0.60 per Mtok input and $3.00 per Mtok output, with a full 256K tokens (262,144) context window.
Atlas with GPT-4.1 mini (2026): The Cheap Slot That Still Holds a Million Tokens
GPT-4.1 mini gives Atlas a 1,047,576 token context at $0.40 per Mtok input and $1.60 per Mtok output. The right small_model when your cheap slot needs a huge window.
Atlas with GPT-OSS 20B (hosted): the cheap slot that can still think in 2026
Run Atlas on GPT-OSS 20B (hosted) in 2026: $0.03/$0.14 per Mtok on DeepInfra, a 131,072 token context, reasoning on, and $0.00/$0.00 locally in LM Studio.
Atlas with Nemotron 70B (Ollama): Instruction Adherence on a Workstation in 2026
Run Atlas on Nemotron 70B (Ollama): NVIDIA's 43GB reward-tuned Llama 3.1 70B with a 128K context, free self-hosted. Setup, memory math, and honest tradeoffs.
Atlas with Databricks Foundation Model APIs in 2026: Frontier Models Inside the Lakehouse
Atlas with Databricks Foundation Model APIs in 2026: Claude Opus 4.7 at a 1,000,000 token context, GPT-5.5 at 1,050,000, governed by the same DATABRICKS_TOKEN.
Atlas with DeepSeek-R1 1.5B Distill (Ollama): The 1.1GB Reasoning Slot in 2026
DeepSeek-R1 1.5B Distill (Ollama) is a 1.1GB reasoning model with a 128K context that runs on CPU. Use it as the Atlas small_model in 2026. Free (self-hosted).
Atlas with StarCoder2 15B (Ollama): The Provenance Choice in 2026
StarCoder2 15B (Ollama) is BigCode's 9.1GB code model with a transparent training corpus and a 16K context, Free (self-hosted). Atlas setup and tradeoffs for 2026.
Atlas with OpenAI o3-pro: The $20 / $80 Reasoning Escape Hatch (2026)
OpenAI o3-pro in Atlas costs $20 per Mtok input and $80 per Mtok output on a 200K context. A one shot escape hatch for hard problems, not an interactive default.
Atlas with GPT-5 Chat: The Only Chat Snapshot That Keeps 400K in 2026
GPT-5 Chat in Atlas: gpt-5-chat-latest is the one chat-latest snapshot that retains a full 400K context and 128K output, at $1.25 per Mtok input.
Atlas with NVIDIA Nemotron Nano 9B v2 in 2026
Nemotron Nano 9B v2 in Atlas, 2026: a dense 9B reasoning model at $0.06/$0.23 per Mtok on Vercel AI Gateway and Amazon Bedrock, free on the NVIDIA NIM tier.
Atlas with OpenCoder 8B (Ollama): the Auditable Local Coder in 2026
OpenCoder 8B (Ollama) runs Atlas on a fully open code LLM: open data, open recipe, 4.7GB, 8K tokens (8,192) of context, Free (self-hosted). Setup and tradeoffs.
Atlas with IBM Granite 4 Small-H (Ollama): a 1M-Token Local Window in 2026
IBM Granite 4 Small-H (Ollama) is a hybrid Mamba model with 1M tokens (1,048,576) of context from a 19GB download, Free (self-hosted). Atlas setup and memory notes.
Atlas with Together AI (gateway) in 2026: Open-Weights Models at Scale
Together AI (gateway) drives Atlas with the broadest open-weights catalog: Qwen3.7 Max at $1.25 / $3.75 per Mtok, up to 1M context, US-hosted inference.
Atlas with GPT-4o in 2026: A 128K Legacy Model with Dated Snapshots
GPT-4o runs in Atlas at $2.50 per Mtok input and $10 per Mtok output on a 128K context with a 16,384 output cap. Best for quick lookups and reproducible baselines.
Atlas with Llama 3.3 70B Instruct (Meta Llama API) in 2026
Llama 3.3 70B Instruct (Meta Llama API) in Atlas for 2026: 128,000 tokens of context, an OpenAI-compatible endpoint, and a 4,096 token output ceiling to plan around.
Atlas with Claude Sonnet 4.6: The Build Agent Default in 2026
Claude Sonnet 4.6 drives Atlas at $3 per Mtok input, $15 per Mtok output on a 1M token window with 128K max output. Setup, cost math, and when Sonnet 5 wins.
Atlas with Azure OpenAI (gateway) in 2026: The GPT-5 Codex Line Under Enterprise IAM
Azure OpenAI (gateway) runs Atlas on the full GPT-5 lineup including gpt-5.3-codex, with up to 1.05M context and passthrough Azure pricing in your own region.
Atlas with Gemma 3 12B Instruct: The Mid-Size Bedrock Gemma for 2026
Gemma 3 12B Instruct in Atlas via Amazon Bedrock: $0.05 per Mtok input, $0.10 per Mtok output, a 131,072 token context, and 12B dense weights you can self-host.
Atlas with GLM-4.7-FlashX: The Cheapest Paid Slot in 2026
GLM-4.7-FlashX runs Atlas at $0.07 per Mtok input and $0.40 per Mtok output on a 200K tokens (200,000) context, with reasoning enabled and a 131,072 output cap.
Atlas with GPT-5 Pro: The 272,000 Token Output Ceiling in 2026
GPT-5 Pro in Atlas: the only OpenAI model with a 272,000 token max output, priced at $15 per Mtok input, $120 per Mtok output on a 400K tokens window.
Atlas with Devstral Small 2: A Free Coding Agent Model in 2026
Devstral Small 2 in Atlas: listed at $0 / $0 per Mtok on Mistral's labs endpoint with a 256K window, or run 24B locally at roughly 14GB via ollama pull devstral.
Atlas with Llama 3.2 3B (local via Ollama): A 2.0GB Offline small_model for 2026
Llama 3.2 3B (local via Ollama) in Atlas for 2026: a 2.0GB pull, Free (self-hosted), 128,000 tokens of context, and the offline small_model that never touches a network.
Atlas with DeepSeek Coder V2 16B Lite (Ollama): The 160K Local Agent Model in 2026
DeepSeek Coder V2 16B Lite (Ollama) gives Atlas a 160K token context from an 8.9GB download, Free (self-hosted). Setup, MoE speed, and the KV cache catch, for 2026.
Atlas with Qwen2.5-Coder 7B (local via Ollama): the Laptop Setup in 2026
Qwen2.5-Coder 7B runs Atlas on a laptop with no discrete GPU: about 5GB at 4-bit, a 32,768 token context, and free self-hosted. Setup, limits, and when to upgrade.
Atlas with Amazon Nova Pro: Bedrock Billing and a 300K Window (2026)
Amazon Nova Pro runs Atlas at $0.80 / $3.20 per Mtok on a 300,000 token window, billed through your existing AWS account with no new vendor contract.
Atlas with Ollama Cloud in 2026: Hosted Ollama Tags for a Terminal Coding Agent
How to run Atlas on Ollama Cloud in 2026: same local tags like qwen3-coder:480b on hosted GPUs, up to 1,048,576 tokens of context, no published per-token price.
Atlas with Qwen3 32B: The Largest Dense Qwen3 in 2026
Qwen3 32B in Atlas: 16,384 max output tokens, hybrid thinking, 128K tokens (131,072) of context, $0.70 per Mtok input and $2.80 per Mtok output in 2026.
Atlas with MiniMax-M2.1 in 2026: A Free Upgrade Over M2
MiniMax-M2.1 runs Atlas at $0.30 per Mtok input and $1.20 per Mtok output with a 204,800 token context and 131,072 max output. Setup, tradeoffs, and when to move on.
Atlas with Amazon Bedrock (gateway) in 2026: One IAM Boundary, Every Model
Amazon Bedrock (gateway) drives Atlas behind AWS IAM: Nova Lite at $0.06 / $0.24 per Mtok, Llama 4 Scout with a 3,500,000 token context, full AWS auth chain.
Atlas with Qwen3-Coder 30B-A3B Instruct: The Default Open Agentic Coder in 2026
Qwen3-Coder 30B-A3B Instruct in Atlas: 256K tokens (262,144), $0.45 per Mtok input, $2.25 per Mtok output, and 3.3B active parameters out of 30B total.
Atlas with Claude Sonnet 5: The 2026 Everyday Driver
Claude Sonnet 5 gives Atlas a 1M token window at $2 / $10 per Mtok, 60 percent under Opus 4.8 input pricing. Why it is the default for long agentic sessions in 2026.
Atlas with Llama 3.3 70B (local via Ollama) in 2026: Dense 70B on Your Own Hardware
Llama 3.3 70B (local via Ollama) drives Atlas at 128K tokens (131,072) and Free (self-hosted) pricing, with a $0.59 / $0.79 per Mtok Groq fallback on identical weights.
Atlas with GPT-5.3 Codex: Code-Specialized Reasoning in 2026
GPT-5.3 Codex is OpenAI's February 2026 code-specialized reasoning model, $1.75 / $14 per Mtok on a 400K window. Built for the long agentic loops Atlas runs.
Atlas with Kimi K2 Thinking Turbo: The 2026 Reasoning Speed Tier
Kimi K2 Thinking Turbo gives Atlas priority serving on a reasoning model at $1.15 per Mtok input and $8.00 per Mtok output, on a 256K tokens (262,144) window.
Atlas with Mistral Medium 3.5: The EU-Hosted Frontier Model (2026)
Mistral Medium 3.5 drives Atlas at $1.50 / $7.50 per Mtok on a 262,144 token window with a matching 262,144 output limit. The EU-hosted option for data residency.
Atlas with GPT-5.6 Terra: The Mid Tier GPT-5.6 Pick for 2026
GPT-5.6 Terra drives Atlas on a 1,050,000 token context at $2.50 per Mtok input, $15 per Mtok output. Setup, cost math, and when Sol or Luna is the better pick.
Atlas with GPT-5.5: The 1.05M Token Jump, Reviewed for 2026
GPT-5.5 took the GPT-5 line to a 1,050,000 token context at $5 per Mtok input, $30 per Mtok output. Atlas setup, the price rise from $1.75, and when GPT-5.6 wins.
Atlas with Claude Opus 4.7 in 2026: 1M Context at $5 / $25 per Mtok
Claude Opus 4.7 drives Atlas with a 1M tokens (1,000,000) window, 128K max output, and $5 per Mtok input, $25 per Mtok output. Setup, tradeoffs, and when to move on.
Atlas with Mixtral 8x22B (local via Ollama): The 80GB Question in 2026
Mixtral 8x22B (local via Ollama) in Atlas for 2026: an 80GB pull, Free (self-hosted), a 64,000 token window, and sparse routing across 8 experts of 22B.
Atlas with Qwen3-Coder 30B (local via Ollama): the Default Local Setup in 2026
Qwen3-Coder 30B is Atlas's default local coding model: a 19GB Ollama download, 256K context, 3.3B active parameters, and free self-hosted. Setup and real tradeoffs.
Atlas with Fireworks AI (gateway) in 2026: Buying Latency with Money
Fireworks AI (gateway) serves Atlas open models with fast-router tiers: DeepSeek V4 Flash at $0.14 / $0.28 per Mtok on a 1,000,000 token context.
Atlas with Grok 4.3: The Cheapest 1M Context Reasoning Model (2026)
Grok 4.3 gives Atlas a 1M token context at $1.25 / $2.50 per Mtok, the cheapest 1M reasoning model from a US lab. Output caps at 30,000 tokens, so plan around it.
Atlas with Kimi K2.7 Code: Open Weights at Trillion-Parameter Scale (2026)
Kimi K2.7 Code drives Atlas at $0.95 / $4 per Mtok on a 262,144 token window. Open weights, 1T total parameters, served by four independent providers.
Atlas with Codestral: Fast Fill-in-the-Middle Editing in the Terminal (2026)
Codestral runs in Atlas at $0.30 / $0.90 per Mtok on a 256K token window. Fast single-file edits, but a 4,096 token output ceiling blocks large refactors.
Atlas with GPT-5 Codex: The First Codex Model of the GPT-5 Family in 2026
GPT-5 Codex in Atlas: September 2025 brought coding post training at zero premium, 400K context, 128K output, and $1.25 per Mtok input with $10 per Mtok output.
Atlas with DeepSeek-V3.1 671B (Ollama): Self-Hosting Frontier Open Weights in 2026
DeepSeek-V3.1 671B (Ollama) is a 404GB MoE with a 160K context and hybrid thinking modes. Run Atlas on frontier open weights behind your own firewall in 2026.
Atlas with IBM Granite Code 8B (Ollama): 125K Context on 4.6GB in 2026
IBM Granite Code 8B (Ollama) gives Atlas a 125K tokens window from a 4.6GB download, Free (self-hosted), with enterprise licensing. Setup, tags, and tradeoffs.
Atlas with GPT-5 Mini: Full 400K Context at One Fifth the Price in 2026
GPT-5 Mini in Atlas: $0.25 per Mtok input and $2 per Mtok output, 5x cheaper than GPT-5 on both sides, with no reduction to the 400K context window.
Atlas with Mistral Medium 3.1 (2508): Setup, Cost, and Fit in 2026
Run Atlas on Mistral Medium 3.1 (2508): a 262,144 token window at $0.40 / 1M input tokens and $2.00 / 1M output tokens. Setup steps, costs, and honest tradeoffs.
Atlas with Mistral 7B: Cost, Context, and Real Limits in 2026
Running Atlas on Mistral 7B in 2026: an 8,000 token window at $0.25 / 1M input tokens. Great for smoke-testing a provider block, wrong for agentic coding.
Atlas with OpenAI o3-mini in 2026: Still Worth Pinning?
OpenAI o3-mini runs in Atlas at $1.10 per Mtok input and $4.40 per Mtok output on a 200K context, but o4-mini costs exactly the same and is newer. Here is the call.
Atlas with OpenAI o3: Cheap Deep Reasoning in the Terminal (2026)
Run Atlas on OpenAI o3 in 2026. A 200K token context reasoning model at $2 per Mtok input and $8 per Mtok output, with real setup steps and honest tradeoffs.
Atlas with DeepSeek Chat: 384,000 Token Output at $0.28 per Mtok in 2026
Run Atlas on DeepSeek Chat in 2026. DeepSeek's non-reasoning endpoint gives 1M tokens (1,000,000) of context at $0.14 per Mtok input, $0.28 per Mtok output.
Atlas with Phi-4 (local via Ollama): The 16K Context Tradeoff in 2026
Phi-4 (local via Ollama) drives Atlas at Free (self-hosted) pricing with a 16K tokens (16,384) window. Setup, the 14B reasoning case, and where it breaks.
Atlas with Gemma 3 27B (local via Ollama): Setup, Cost, and Tradeoffs in 2026
Run Atlas on Gemma 3 27B (local via Ollama) in 2026: a 131,072 token context, Free (self-hosted) pricing, single-GPU inference, and the honest tradeoffs.
Atlas with Gemini 3.1 Pro: Setup, Cost, and Tradeoffs in 2026
Run Atlas on Gemini 3.1 Pro in 2026: a 1,048,576 token window at $2 / $12 per Mtok, with real setup steps, the 65,536 output ceiling, and when to switch models.
Atlas with Amazon Nova 2 Lite: Cost, Context, and Setup in 2026
Amazon Nova 2 Lite drives Atlas from inside your AWS account at $0.33 / $2.75 per Mtok with a 128K context. Setup, real tradeoffs, and when to pick another model.
Atlas with Gemini 3.5 Flash: The May 2026 Latency Tier That Keeps 1M Context
Gemini 3.5 Flash in Atlas: the May 2026 release keeps the full 1,048,576 token window at $1.50 / $9 per Mtok, with reasoning and tool calling enabled.
Atlas with Qwen2.5-Coder 1.5B (Ollama): the 986MB small_model slot in 2026
Qwen2.5-Coder 1.5B (Ollama) in Atlas: a 986MB Q4_K_M pull with 32K tokens (32,768) of context, Free (self-hosted), sized for titles and summaries, not refactors.
Atlas with Claude Opus 4.8: The 2026 Model Guide
Claude Opus 4.8 in Atlas: a 1M token context window at $5 / $25 per Mtok with 128K output. When to pin Anthropic's top coding model in 2026, and when not to.
Atlas with DeepSeek V4 Flash: The Cheapest 1M Context Reasoning Model in 2026
DeepSeek V4 Flash in Atlas: $0.14 / $0.28 per Mtok on a 1M window with 384,000 output tokens, or $0.09 / $0.18 via DeepInfra. Setup, small_model wiring, tradeoffs.
Atlas with Mistral Small 3.2 (2506): The Cheap Slot Done Right in 2026
Mistral Small 3.2 (2506) gives Atlas a 128,000 token window at $0.10 / 1M input tokens and $0.30 / 1M output tokens. Setup, the 16,384 token output cap, tradeoffs.
Atlas with Gemini 2.5 Pro: The Cheap 1M Context Default in 2026
Gemini 2.5 Pro in Atlas: a stable GA id with a 1,048,576 token window at $1.25 per Mtok input, 37 percent cheaper to read with than Gemini 3 Pro.
Atlas with GPT-5.1 Codex mini: Wide Subagent Fan Out in 2026
GPT-5.1 Codex mini in Atlas: an OpenAI fast tier at $0.25 per Mtok input, $2 per Mtok output, built for parallel subagents, bulk sweeps, and cheap summaries.
Atlas with Magistral Small (local via Ollama): Private Reasoning in 2026
Magistral Small (local via Ollama) in Atlas for 2026: a 14GB pull, Free (self-hosted), reasoning traces generated on your own GPU, and an honest 40,000 token cap.
Atlas with Command R: The $0.15 per Mtok Background Model for 2026
Command R runs Atlas background work at $0.15 per Mtok input and $0.6 per Mtok output, with a 128,000 token context and a 4,000 token output cap on patches.
Atlas with Qwen3 8B: The Cheap Thinking Model for small_model in 2026
Qwen3 8B in Atlas for 2026: hybrid thinking at $0.18 per Mtok input and $0.70 per Mtok output, 128K tokens (131,072), and a natural small_model slot fit.
Atlas with OpenAI o1 in 2026: The Original Reasoning Model, Priced Like It
OpenAI o1 runs in Atlas at $15 per Mtok input and $60 per Mtok output on a 200K context. Historically important, now dominated by o3 on both price and capability.
Atlas with QwQ Plus: Alibaba's Reasoning Line in the Plan Agent, 2026
Run Atlas on QwQ Plus in 2026. Alibaba's dedicated reasoning tier costs $0.80 per Mtok input, $2.40 per Mtok output, with a 128K tokens (131,072) context.
Atlas with Gemini Flash-Lite Latest: A Set-and-Forget small_model for 2026
gemini-flash-lite-latest in Atlas: a rolling alias for Google's cheapest reasoning-capable Lite tier at $0.1 per Mtok input and a 1,048,576 token context.
Atlas with Vercel AI Gateway in 2026: 310 Models Behind One AI_GATEWAY_API_KEY
Atlas with Vercel AI Gateway in 2026: roughly 310 models, Grok 4.20 Reasoning at a 2,000,000 token context for $1.25/$2.50 per Mtok, one AI_GATEWAY_API_KEY.
Atlas with GLM-5.1: Reasoning, Cost, and Context in 2026
GLM-5.1 from Z.ai drives Atlas with a 200,000 token context at $1.40 per Mtok input and $4.40 per Mtok output. Setup, honest tradeoffs, and how it compares to GLM-5.2.
Atlas with Qwen3-Coder 480B (Ollama): the self-hosted ceiling in 2026
Qwen3-Coder 480B (Ollama) in Atlas: 290GB of weights, roughly 292GB to serve, 256K tokens (262,144) of context. Free (self-hosted), but the hardware is not.
Atlas with IBM Granite 4.1 8B in 2026
IBM Granite 4.1 8B in Atlas, 2026: a dense 8B model at $0.05/$0.10 per Mtok with a 131,072 token context where max output equals the full window.
Atlas with Snowflake Cortex in 2026: Claude Opus 4.8 Inside Your Snowflake Boundary
Atlas with Snowflake Cortex in 2026: Claude Opus 4.8 and Claude Fable 5 at a 1,000,000 token context, governed by Snowflake RBAC and billed in Snowflake credits.
Atlas with Code Llama 34B (Ollama): The Practical Top of the Line in 2026
Code Llama 34B (Ollama) is 19GB, roughly 21GB to serve, with a 16K context and Free (self-hosted) pricing. Atlas setup, why 34B is the last sane size, for 2026.
Atlas with Llama 3.3 8B Instruct (Meta Llama API): The small_model Slot in 2026
Llama 3.3 8B Instruct (Meta Llama API) in Atlas for 2026: 128,000 tokens of context in an 8B-class model, a 4,096 token output cap, and why it belongs in small_model.
Atlas with Amazon Nova Lite in 2026: 300K Context on Your AWS Bill
Amazon Nova Lite drives Atlas at $0.06 per Mtok input and $0.24 per Mtok output with a 300K token context, billed through IAM with no new vendor API key to manage.
Atlas with Mistral Small 3.2 (local via Ollama) in 2026
Run Atlas on Mistral Small 3.2 (local via Ollama) in 2026: a 15GB pull, Free (self-hosted), 128,000 tokens of context, and function-calling tuned 2506 weights.
Atlas with Code Llama 7B (Ollama): A 3.8GB Local Coder in 2026
Code Llama 7B (Ollama) is Meta's 2023 code model at 3.8GB with a 16K context, Free (self-hosted). Atlas setup, the :code and :python tags, and 2026 tradeoffs.
Atlas with Ministral 3B: The $0.04 Housekeeping Model in 2026
Ministral 3B is the cheapest model Mistral sells: $0.04 / 1M input tokens and $0.04 / 1M output tokens across 128,000 tokens. Atlas small_model setup and limits.
Atlas with Hugging Face Inference in 2026: 51 Models Behind One HF_TOKEN
Atlas with Hugging Face Inference in 2026: 51 routed models behind one HF_TOKEN, GLM-4.7-Flash free at $0/$0 per Mtok, and the router margin on popular rows.
Atlas with GPT-5.4 Pro: One Shot Hard Problems in 2026
GPT-5.4 Pro is the max reasoning tier at $30 per Mtok input, $180 per Mtok output on a 1,050,000 token window. Why Atlas users switch to it instead of pinning it.
Atlas with QwQ 32B (Ollama): a free local reasoning model for the plan agent in 2026
QwQ 32B (Ollama) in Atlas: Qwen's dedicated reasoning model at 20GB, 40K tokens (40,960) of context, Free (self-hosted). Let QwQ plan, then hand off to a coder.
Atlas with Google Vertex AI (gateway) in 2026: Gemini and Claude Under One GCP Project
Google Vertex AI (gateway) runs Atlas on Gemini 3.1 Pro at $2 / $12 per Mtok with a 1M context, plus Claude through Atlas's google-vertex-anthropic route.
Atlas with Qwen3.6 35B-A3B: The $0.248 per Mtok Workhorse of 2026
Qwen3.6 35B-A3B is the cheapest reasoning model in the Qwen3.6 line at $0.248 per Mtok input and $1.485 per Mtok output, with a full 256K tokens (262,144) context for Atlas.
Atlas with Qwen3 4B (Ollama): 256K of context on a 2.5GB pull in 2026
Qwen3 4B (Ollama) in Atlas: a 2.5GB download carrying 256K tokens (262,144) of context, Free (self-hosted), with separate 2507 instruct and thinking tags.
Atlas with GPT-5 Nano: The Cheapest Model in the OpenAI Registry in 2026
GPT-5 Nano in Atlas: $0.05 per Mtok input and $0.40 per Mtok output, the cheapest model in the OpenAI registry, and the right pick for the small_model slot.
Atlas with Qwen2.5 72B Instruct: The Flagship Dense Qwen in 2026
Qwen2.5 72B Instruct in Atlas: 128K tokens (131,072), $1.40 per Mtok input, $5.60 per Mtok output, openly published weights you can serve on your own vLLM.
Atlas with DeepInfra: The Cheapest Open-Weights Host for an Agent Loop in 2026
Run Atlas on DeepInfra: GPT OSS 120B at $0.037/$0.17 per Mtok, DeepSeek V4 Flash at a 1,048,576 token window for $0.09 input. Setup, limits, and cost math.
Atlas with Grok 4.20 Multi-Agent in 2026: Two Levels of Fan-Out
Grok 4.20 Multi-Agent orchestrates internal agents behind one model id, stacking with Atlas's own parallel subagents. 1,000,000 token context at $1.25 / $2.5 per Mtok.
Atlas with Magistral Small: Open Reasoning for the Plan Agent in 2026
Magistral Small is Mistral's first open reasoning model: 128,000 tokens at $0.50 / 1M input tokens and $1.50 / 1M output tokens. Atlas setup, costs, and tradeoffs.
Atlas with Kimi K2 Thinking: Open-Weights Reasoning at $0.60 per Mtok (2026)
Kimi K2 Thinking gives Atlas reasoning at $0.60 / $2.50 per Mtok on a 262,144 token window. Open weights, and it sustains the long tool-calling chains agents need.
Atlas with Kimi K2 0905: 262,144 Tokens In and Out, 2026
Run Atlas on Kimi K2 0905 in 2026. Moonshot's September refresh gives 256K tokens (262,144) context and output at $0.60 per Mtok in, $2.50 per Mtok out.
Atlas with Qwen Flash: A Cheap Tier That Still Writes Real Diffs in 2026
Run Atlas on Qwen Flash in 2026. Alibaba's fast tier pairs 1M tokens (1,000,000) of context with a 32,768 token output at $0.05 per Mtok in, $0.40 per Mtok out.
Atlas with Code Llama (local via Ollama): a Fill-in-the-Middle Baseline in 2026
Code Llama runs free and local via Ollama at 7b, 13b, 34b, and 70b with a 16,384 token context. Strong at infilling, but too small and too old for Atlas agent work.
Atlas with Qwen3.6 Flash: A $0.1875 Speed Tier With Coding Lineage in 2026
Qwen3.6 Flash in Atlas: Alibaba's April 2026 speed tier at $0.1875 / $1.125 per Mtok with a 1M window, wired to small_model in atlas.json. Setup and tradeoffs.
Atlas with Qwen3 Coder Plus: Agentic Coding on a 1M Window in 2026
Qwen3 Coder Plus in Atlas: a 1,048,576 token window at $1 / $5 per Mtok, post-trained for agentic coding, with ollama pull qwen3-coder:30b as the local counterpart.
Atlas with Llama 4 Scout: the 3.5M Token Context Model in 2026
Llama 4 Scout gives Atlas a 3.5M token context on Bedrock at $0.17 / $0.66 per Mtok, or $0.10 / $0.30 on DeepInfra. Setup, the portability trap, and when to switch.
Atlas with GPT-5.4: The Balanced 2026 Default
GPT-5.4 brings a 1,050,000 token window to Atlas at $2.50 / $15 per Mtok, half the input cost of GPT-5.6. The balanced day-to-day default for 2026 sessions.
Atlas with Gemma 3 12B (Ollama): 128K Context and Screenshots in 2026
Gemma 3 12B (Ollama) gives Atlas a 128K tokens (131,072) window and multimodal image input from an 8.1GB download, Free (self-hosted). Setup, sizing, and limits.
Atlas with DeepSeek Coder 33B (Ollama): Local Setup and Tradeoffs in 2026
DeepSeek Coder 33B (Ollama) drives Atlas locally from 19GB of weights with a 16K token context and no API bill. Setup, VRAM budget, and honest tradeoffs for 2026.
Atlas with AI21 Jamba Large 1.7 in 2026
AI21 Jamba Large 1.7 in Atlas, 2026: a hybrid SSM-Transformer with 256,000 tokens of context at $2.00/$8.00 per Mtok, capped at 4,096 tokens of output.
Atlas with Codestral 22B (Ollama): 32K Context and a License Gate in 2026
Codestral 22B (Ollama) is Mistral's 13GB code model with a 32K context, fluent across many languages. Free (self-hosted), non-commercial license. Atlas setup for 2026.
Atlas with Upstage Solar Pro 3 in 2026
Solar Pro 3 in Atlas, 2026: symmetric $0.25/$0.25 per Mtok on the Upstage API at 131,072 tokens, versus $0.15/$0.60 and 128,000 tokens on OpenRouter.
Atlas with Mistral Medium 3 (2505) in 2026: Symmetric Limits, $0.40 In
Mistral Medium 3 (2505) runs Atlas with symmetric 131,072 token context and output at $0.40 / 1M input tokens and $2.00 / 1M output tokens. Setup, limits, successors.
Atlas with Magistral Medium: Multi-Hop Root-Cause Debugging in 2026
Magistral Medium reasons across a 128,000 token window at $2.00 / 1M input tokens and $5.00 / 1M output tokens. Atlas setup, the 16,384 token output cap, tradeoffs.
Atlas with Cerebras (gateway) in 2026: Wafer-Scale Speed, Three Models
Cerebras (gateway) drives Atlas on wafer-scale engines: GPT-OSS 120B at $0.35 / $0.75 per Mtok, 131K tokens (131,072) context, and a catalog of three models.
Atlas with Gemini 3.1 Flash Lite: The $0.25 Small Model Slot in 2026
Gemini 3.1 Flash Lite in Atlas: 1,048,576 tokens of context at $0.25 / $1.50 per Mtok, wired to the small_model slot in atlas.json. Setup, limits, and tradeoffs.
Atlas with DeepSeek Reasoner: Chain-of-Thought Debugging at $0.28 in 2026
DeepSeek Reasoner in Atlas: visible reasoning traces on a 1M window at $0.14 / $0.28 per Mtok, roughly 640x below GPT-5.5 Pro's $180 output rate. Setup and limits.
Atlas with Claude Fable 5: Plan Mode Model Guide for 2026
Claude Fable 5 is Anthropic's premium June 2026 model at $10 / $50 per Mtok with a 1M window. Use it for one expensive Atlas planning pass, then drop back down.
Atlas with GPT-5.1 Codex Max: The Cheapest Codex Tier in 2026
GPT-5.1 Codex Max still bills at the November 2025 rate of $1.25 / $10 per Mtok on a 400K window. The lowest-cost entry into OpenAI's Codex post-training for Atlas.
Atlas with GPT-5.1: The Cheapest Full Size GPT-5 Input Price in 2026
GPT-5.1 in Atlas: 400K context, $1.25 per Mtok input and $10 per Mtok output, cheaper on input than every GPT-5 release that followed it in 2026.
Atlas with Qwen2.5 14B Instruct in 2026: The Open-Weights Main Model Tier
Qwen2.5 14B Instruct drives Atlas at $0.35 per Mtok input and $1.40 per Mtok output, fits a single 24 GB GPU at 4-bit, and holds a real multi-file edit plan.
Atlas with GLM-4.5-Air: The 106B Self-Hostable Cheap Slot in 2026
GLM-4.5-Air drives Atlas at $0.20 per Mtok input and $1.10 per Mtok output on a 128K tokens (131,072) window. A 106B total / 12B active MIT-licensed MoE.
Atlas with DeepSeek R1 (0528): The Open Reasoning Trace, 2026
Run Atlas on DeepSeek R1 (0528) in 2026. DeepInfra hosts the MIT-licensed open reasoning model at $0.50 per Mtok in, $2.15 per Mtok out, 160K tokens context.
Atlas with Qwen3 235B-A22B: Flagship Sparse Reasoning in 2026
Qwen3 235B-A22B in Atlas: 235B total parameters, 22B active per token, $0.70 per Mtok input and $2.80 per Mtok output, 128K tokens (131,072) of context.
Atlas with Qwen2.5 72B (local via Ollama): Air-Gapped Coding in 2026
Run Atlas fully offline on Qwen2.5 72B (local via Ollama) in 2026. Free (self-hosted), about 47 GB at Q4_K_M, and a 32,768 token local context cap.
Atlas with Gemini 2.5 Flash: The Workhorse small_model Pick for 2026
Gemini 2.5 Flash in Atlas: reasoning enabled, a 1,048,576 token context, and $0.3 per Mtok input, a sixth of what a frontier tier model charges to read the same repo.
Atlas with GPT-OSS 120B (hosted): picking the right host in 2026
Run Atlas on GPT-OSS 120B (hosted) in 2026. Identical Apache weights cost $0.037 per Mtok on DeepInfra and $0.35 on Cerebras, a 9.5x input spread. Full host guide.
Atlas with GPT-5.4 nano: The $0.20 Background Model in 2026
GPT-5.4 nano is OpenAI's cheapest reasoning-capable model at $0.20 / $1.25 per Mtok with a 400K window. Built for the titles, summaries, and classification Atlas runs constantly.
Atlas with Qwen Turbo: The $0.05 per Mtok Small Model Slot in 2026
Run Atlas on Qwen Turbo in 2026. Alibaba's cheapest reasoning tier gives 1M tokens (1,000,000) of context at $0.05 per Mtok input, $0.20 per Mtok output.
Atlas with Qwen3.5 27B: The Dense Entry Point to Qwen3.5 in 2026
Qwen3.5 27B gives Atlas 256K tokens (262,144) of context and predictable dense latency at $0.30 per Mtok input and $2.40 per Mtok output. Setup, costs, and honest tradeoffs.
Atlas with StarCoder2 (local via Ollama): Auditable Training Data in 2026
StarCoder2 is BigCode's open code model, trained on The Stack v2 with full data provenance. Free self-hosted, 600-plus languages, 16,384 token context. Atlas setup.
Atlas with Perplexity Sonar in 2026: The Live-Web Research Model, Not the Build Model
Atlas with Perplexity Sonar in 2026: live web results with citations, Sonar at $1.00/$1.00 per Mtok, Sonar Pro at 200,000 tokens, and why it is not a build model.
Atlas with Gemma 4 12B (Ollama): 256K Context from a 7.6GB Download in 2026
Gemma 4 12B (Ollama) is a 7.6GB download with a 256K tokens (262,144) context, Free (self-hosted). The longest Gemma window that fits a mid-range GPU. Atlas setup.
Atlas with Poolside Laguna M.1 in 2026
Poolside Laguna M.1 in Atlas, 2026: a coding-native reasoning model with 262,144 tokens of context, free on Poolside's API and $0.20/$0.40 per Mtok on OpenRouter.
Atlas with Devstral Small 2505: The Original Agent-First Model in 2026
Devstral Small 2505 started Mistral's Devstral line in May 2025: 128,000 tokens at $0.10 / 1M input tokens. Atlas setup, why it was superseded, and when to pin it.
Atlas with OpenAI o4-mini (2026): Cheap Reasoning for Parallel Subagents
OpenAI o4-mini drives Atlas at $1.10 per Mtok input and $4.40 per Mtok output on a 200K context. A strict upgrade over o3-mini at identical price, with real tradeoffs.
Atlas with Qwen Plus: 1M Context for $0.40 per Mtok in 2026
Run Atlas on Qwen Plus in 2026. Alibaba's mid tier gives 1M tokens (1,000,000) of context with reasoning at $0.40 per Mtok input, $1.20 per Mtok output.
Atlas with Qwen3.5 397B-A17B: The Qwen3.5 Flagship in 2026
Qwen3.5 397B-A17B is Alibaba's Qwen3.5 flagship: 397B total, 17B active, 256K tokens (262,144) of context, $0.60 per Mtok input and $3.60 per Mtok output, running in Atlas.
Atlas with Command R7B in 2026: Cohere's Cheapest Model, Used Correctly
Command R7B costs $0.0375 per Mtok input, about 1/66th of Command A, and still carries a 128,000 token context. Here is how to slot it into Atlas without wrecking your code.
Atlas with Mixtral 8x22B: The Largest Open MoE of Its Era in 2026
Mixtral 8x22B scaled the MoE idea in April 2024: 64,000 tokens at $2.00 / 1M input tokens and $6.00 / 1M output tokens. Atlas setup, self-hosting, and honest limits.
Atlas with Claude Opus 4.1: The Pre Price Cut Opus in 2026
Claude Opus 4.1 runs Atlas on a 200K context at $15 per Mtok input, $75 per Mtok output, with a 32K output ceiling. Setup, cost warnings, and better alternatives.
Atlas with Gemini 2.0 Flash: A Fast Reader With an 8,192 Token Output Cap in 2026
Gemini 2.0 Flash in Atlas: the December 2024 release with a 1,048,576 token context, $0.1 per Mtok input, no reasoning mode, and an 8,192 token output ceiling.
Atlas with Grok Build 0.1: xAI's First Coding-Agent Model (2026)
Grok Build 0.1 runs in Atlas at $1 / $2 per Mtok with a 256K context and 256,000 max output tokens. A coding-specialized model, but it is a 0.1 release.
Atlas with GPT-5.4 mini: A Reasoning Mini Tier for 2026
GPT-5.4 mini runs $0.75 / $4.50 per Mtok on a 400K window, undercutting Claude Haiku 4.5's $1 with double the context. A reasoning-capable small_model for Atlas.
Atlas with Gemma 4 26B A4B: The Open-Weights MoE Option in 2026
Gemma 4 26B A4B in Atlas: a sparse MoE with about 4B active of 26B total parameters, a 262,144 token context, a 32,768 token output cap, and no public price.
Atlas with Gemini 3.1 Pro Custom Tools: Cost, Context, and Setup in 2026
Run Atlas on Gemini 3.1 Pro Custom Tools in 2026: a 1,048,576 token context, $2 per Mtok input, and a tool-calling checkpoint built for a dense tool surface.
Atlas with MiniMax-M2.5 in 2026: Reasoning Under a Dollar
MiniMax-M2.5 drives Atlas on a 230B efficient-MoE at $0.30 per Mtok input and $1.20 per Mtok output, with a 204,800 token context and 131,072 max output tokens.
Atlas with SiliconFlow in 2026: Qwen3 Coder 480B at $0.25/$1.00 per Mtok
Atlas with SiliconFlow in 2026: Qwen3-Coder-480B-A35B at $0.25/$1.00 per Mtok, a free Qwen3.5-4B small_model, and the api.siliconflow.cn residency question.
Atlas with Llama 4 Scout (Ollama): A 10M-Token Window on Local Hardware in 2026
Run Atlas on Llama 4 Scout (Ollama): a 67GB 16-expert MoE with a 10M-token context and image input, free self-hosted. The memory math before you pull 67GB.
Atlas with DeepCoder 14B (Ollama): RL-Tuned for First-Attempt Diffs in 2026
DeepCoder 14B (Ollama) is a 9.0GB RL-tuned coder with a 128K context, built for first-attempt correctness. Free (self-hosted). Atlas setup and tradeoffs for 2026.
Atlas with NVIDIA Nemotron 3 Ultra 550B A55B in 2026
Run Atlas on NVIDIA Nemotron 3 Ultra 550B A55B in 2026. A 550B mixture of experts with 55B active, 1,000,000 tokens on NVIDIA NIM at $0.50/$2.50 per Mtok.
Atlas with Phi-4 Mini 3.8B (Ollama): Native Function Calling at 2.5GB in 2026
Phi-4 Mini 3.8B (Ollama) is a 2.5GB model with 128K tokens (131,072) of context and native function calling, Free (self-hosted). The Atlas small_model that can route tools.
Atlas with MiniMax-M2: The $0.30 Open-Weights Agent Model in 2026
MiniMax-M2 runs Atlas at $0.30 per Mtok input and $1.20 per Mtok output with a 196,608 token context, open weights on HuggingFace, and an Anthropic-compatible API.
Atlas with Mistral Small 24B (Ollama): The Single-GPU Commercial Pick for 2026
Run Atlas on Mistral Small 24B (Ollama): 14GB weights, 32K context, Apache 2.0, free self-hosted. Setup, the 32K vs 128K tag trap, and when a coder beats it.
Atlas with Liquid AI LFM2-24B-A2B in 2026
Liquid AI LFM2-24B-A2B in Atlas, 2026: a liquid neural network MoE at $0.03/$0.12 per Mtok on Together AI, with a 32,768 token context and matching output.
Atlas with GitHub Models in 2026: Free Model Access Behind a GITHUB_TOKEN
Run Atlas on GitHub Models in 2026: every model listed at $0/$0 per Mtok, auth with the GITHUB_TOKEN you already have, and 256,000 tokens on AI21 Jamba 1.5 Large.
Atlas with Baseten in 2026: 262,000 Tokens In, 262,000 Tokens Out
Atlas with Baseten in 2026: Kimi K2.7 Code at 262,000 tokens in and 262,000 out, GPT OSS 120B at $0.10/$0.50 per Mtok, and the tradeoffs of dedicated hosting.
Atlas with Llama 3.2 3B (Ollama): The CPU-Only Floor for a Local Agent in 2026
Run Atlas on Llama 3.2 3B (Ollama): a 2.0GB model with a 128K context, free self-hosted. The realistic floor for an Atlas setup with no GPU at all.
Atlas with IBM Granite 4.0 H Micro in 2026
IBM Granite 4.0 H Micro in Atlas, 2026: a hybrid Mamba-Transformer model at $0.017/$0.112 per Mtok on Cloudflare Workers AI, the cheapest input in the registry.
Atlas with Mistral Nemo 12B (local via Ollama): The 12GB GPU Pick for 2026
Mistral Nemo 12B (local via Ollama) in Atlas for 2026: a 7.1GB pull, Free (self-hosted), Tekken tokenizer, and why the KV cache, not the weights, caps context.
Atlas with Mistral Large 3 (2512): Big Diffs, EU Hosted, 2026
Mistral Large 3 (2512) drives Atlas with a 262,144 token context and a matching 262,144 token output at $0.50 / 1M input tokens and $1.50 / 1M output tokens.
Atlas with CodeGemma 7B (Ollama): Fill-in-the-Middle on 8GB in 2026
CodeGemma 7B (Ollama) is Google's 5.0GB code model with fill-in-the-middle training and an 8K context, Free (self-hosted). Atlas setup and honest tradeoffs for 2026.
Atlas with IBM Granite 3.3 8B (Ollama): the Free Local small_model for 2026
IBM Granite 3.3 8B (Ollama) is a 4.9GB general model with 128K tokens (131,072) of context, Free (self-hosted). Assign it to small_model in Atlas. Setup and limits.
Atlas with GLM-5-Turbo: Setup, Pricing, and Tradeoffs in 2026
Run Atlas on GLM-5-Turbo from Z.ai. A 200,000 token context at $1.20 per Mtok input and $4.00 per Mtok output, plus why Turbo costs more than base GLM-5.
Atlas with Mistral Large 2.1 (2411): A 2026 Setup Guide
Mistral Large 2.1 (2411) runs Atlas on EU infrastructure with a 131,072 token context at $2.00 / 1M input tokens and $6.00 / 1M output tokens. Setup and honest limits.
Atlas with Gemini Flash Latest: The Rolling Alias Explained (2026)
gemini-flash-latest in Atlas: a rolling alias, not a pinned checkpoint. $0.3 per Mtok input, $2.5 per Mtok output, 1,048,576 token context, and no reproducibility.
Atlas with Kimi K2 0711: The Original Trillion-Parameter Preview in 2026
Run Atlas on Kimi K2 0711 in 2026. Moonshot's original K2 preview costs $0.60 per Mtok input, $2.50 per Mtok output, with a 128K tokens (131,072) context.
Atlas with Qwen3 32B (Ollama): dense reasoning over raw speed in 2026
Qwen3 32B (Ollama) in Atlas: dense 20GB weights on a 24GB card, 40K tokens (40,960) of context, Free (self-hosted). Slower than the MoE, steadier on hard problems.
Atlas with Qwen3-Next 80B-A3B Thinking: The Reasoning Tier in 2026
Drive Atlas with Qwen3-Next 80B-A3B Thinking: a reasoning trace over 128K tokens (131,072) of context at $0.50 per Mtok input and $6.00 per Mtok output.
Atlas with Qwen3.6 Plus: The Stable Million-Token Tier in 2026
Qwen3.6 Plus gives Atlas 1M tokens (1,000,000) of context at $0.50 per Mtok input and $3.00 per Mtok output, matching Qwen3.7 Plus while keeping 3.6 generation behavior.
Atlas with Kimi K2 Turbo: Pricing, Context, and Setup in 2026
Kimi K2 Turbo in Atlas costs $2.40 per Mtok input and $10.00 per Mtok output on a 256K tokens (262,144) window. A latency purchase, not a capability upgrade.
Atlas with DeepSeek V3 (open weights): The Frozen 671B Baseline in 2026
Run Atlas on DeepSeek V3 (open weights) in 2026. DeepInfra hosts the MIT-licensed 671B MoE at $0.32 per Mtok input, $0.89 per Mtok output, 128K context.
Atlas with DeepSeek V3.1 (open weights): Togglable Thinking in 2026
Run Atlas on DeepSeek V3.1 (open weights) in 2026. One MIT-licensed checkpoint with thinking and non-thinking modes, $0.25 per Mtok in and $0.95 per Mtok out.
Atlas with GLM-5.2: A 1M Token Open-Weights Model at $1.40 per Mtok (2026)
GLM-5.2 drives Atlas on a 1M token context at $1.40 / $4.40 per Mtok. Z.ai's June 2026 flagship, the first GLM to reach 1M, with open-weights lineage.
Atlas with Qwen3-Coder Next (local via Ollama): the Top-Ranked Local Coder in 2026
Qwen3-Coder Next is the top-ranked local coding model of mid-2026: 262,144 token context, free self-hosted, or $0.22 / $1.80 per Mtok on Bedrock. Atlas setup and tradeoffs.
Atlas with Upstage Solar in 2026
Run Atlas on the Upstage Solar lineup in 2026: solar-pro3 and solar-pro2 at $0.25/$0.25 per Mtok, solar-mini at $0.15/$0.15, across three context sizes.
Atlas with Qwen3 30B-A3B (Ollama): the MoE throughput trade in 2026
Qwen3 30B-A3B (Ollama) in Atlas: 30B total parameters, roughly 3B active per token, 19GB of weights, 256K tokens (262,144) of context, Free (self-hosted).
Atlas with Llama 3.1 8B (Ollama): 128K Context on an 8GB Card in 2026
Run Atlas on Llama 3.1 8B (Ollama): Meta's 4.9GB workhorse with a 128K context, free self-hosted. Setup, the KV cache catch, and when a coder model wins.
Atlas with Gemma 2 27B (Ollama): the Free Local Diff Reviewer in 2026
Gemma 2 27B (Ollama) is Google's 2024 flagship open model, 16GB and Free (self-hosted), with 8K tokens (8,192) of context. Use it to review Atlas diffs, not write them.
Atlas with Phi-3 Medium 14B (Ollama): 128K Context at 7.9GB in 2026
Phi-3 Medium 14B (Ollama) gives Atlas a 128K tokens (131,072) window from a 7.9GB download, Free (self-hosted). Pull phi3:14b, not :medium-4k. Setup and tradeoffs.
Atlas with Gemma 4 31B (Ollama): the Flagship Gemma 4 Tag in 2026
Gemma 4 31B (Ollama) is the largest Gemma 4 tag: 20GB weights, a 256K tokens (262,144) context, Free (self-hosted), plus a 31b-coding-mtp-bf16 variant. Atlas setup.
Atlas with Qwen2.5-Coder 7B (Ollama): the default fully offline setup for 2026
Qwen2.5-Coder 7B (Ollama) in Atlas: the 4.7GB default tag with 32K tokens (32,768) of context, Free (self-hosted), and no cloud key anywhere in the agent loop.
Atlas with Gemma 3 27B Instruct on Amazon Bedrock: Flat Pricing, 202,752 Tokens (2026)
Gemma 3 27B Instruct in Atlas via Amazon Bedrock: $0.12 per Mtok input, $0.2 per Mtok output, a 202,752 token context, open weights, and an 8,192 token output cap.
Atlas with Gemma 3 4B Instruct: A Triage Model, Not a Builder (2026)
Gemma 3 4B Instruct in Atlas via Amazon Bedrock: $0.04 per Mtok input, $0.08 per Mtok output, a 128K context, and a 4,096 token output cap that rules out diffs.
Atlas with GPT-5: The Original 400K Reasoning Model in 2026
GPT-5 in Atlas: the August 2025 launch model with a 400K context, 128K max output, and $1.25 per Mtok input, still the floor price for a full size GPT-5 class model.
Atlas with Command A: Cohere's 256K Context Flagship in 2026
Command A gives Atlas a 256,000 token read window at $2.5 per Mtok input and $10 per Mtok output, with an 8,000 token output cap that shapes how you refactor.
Atlas with North Mini Code in 2026: A 64,000 Token Output Budget
North Mini Code gives Atlas a 256,000 token context and a 64,000 token output, 8x Command A's 8,000, listed at $0 per Mtok on both input and output in the registry.
Atlas with Command A Reasoning: Reasoning You Can Deploy On-Prem (2026)
Command A Reasoning gives Atlas a 256K window at $2.50 / $10 per Mtok, and Cohere lets you run it on-prem or in a VPC. Reasoning with no price premium.
Atlas with Groq (gateway) in 2026: LPU Speed for the Agent Loop
Groq (gateway) runs open models on LPU hardware for Atlas: GPT-OSS 120B at $0.15 / $0.60 per Mtok, 131K tokens (131,072) context, and no Claude or GPT-5.
Atlas with Qwen2.5 7B Instruct in 2026: The Cheap Dense Small Model
Qwen2.5 7B Instruct runs Atlas's small_model slot at $0.175 per Mtok input and $0.70 per Mtok output, keeping the full 131,072 token window on a dense 7B checkpoint.
Atlas with Kimi K2.6: The Generalist Reasoning Tier in 2026
Kimi K2.6 drives Atlas at $0.95 per Mtok input and $4.00 per Mtok output on a 256K tokens (262,144) window. The generalist pick when the job is not purely code.
Atlas with Kimi K2.7 Code HighSpeed: The 2x Throughput Tier in 2026
Kimi K2.7 Code HighSpeed runs Atlas at $1.90 per Mtok input and $8.00 per Mtok output, exactly 2x K2.7 Code, on the same 256K tokens (262,144) context window.
Atlas with Qwen3-Coder 480B-A35B Instruct: The Open Frontier Coder in 2026
Qwen3-Coder 480B-A35B Instruct in Atlas: 480B total parameters, 35B active per token, 262,144 tokens of context, $1.50 per Mtok in and $7.50 per Mtok out.
Atlas with Mistral Small 4 (2603): Cheap Reasoning in 2026
Mistral Small 4 (2603) brings reasoning to the Small tier: 256,000 tokens at $0.15 / 1M input tokens and $0.60 / 1M output tokens. Atlas setup, costs, tradeoffs.
Atlas with DeepSeek V3.2 (open weights): Sparse Attention at $0.38 Output, 2026
Run Atlas on DeepSeek V3.2 (open weights) in 2026. DeepSeek Sparse Attention gives 160K tokens (DeepInfra) at $0.26 per Mtok in and $0.38 per Mtok out.
Atlas with Grok 4.20 (Reasoning) in 2026: A 1M Token Reader
Grok 4.20 (Reasoning) reads 1,000,000 tokens at $1.25 per Mtok input and writes at $2.5 per Mtok, but caps output at 30,000 tokens. A superb reader, a terse writer.
Atlas with Poolside Laguna XS 2.1 in 2026
Poolside Laguna XS 2.1 in Atlas, 2026: the fast tier of Poolside's coding-native line at $0.06/$0.12 per Mtok on OpenRouter, holding 262,144 tokens of context.
Atlas with NVIDIA Nemotron 3 Nano 30B A3B in 2026
Nemotron 3 Nano 30B A3B in Atlas, 2026: 3B active parameters at $0.05/$0.20 per Mtok on DeepInfra, free on NVIDIA NIM, up to 1,048,576 tokens on Ollama Cloud.
Atlas with Command R 35B (Ollama): A RAG-Native Model for Retrieval-Heavy Work in 2026
Run Atlas on Command R 35B (Ollama): Cohere's 19GB RAG and tool-use model with a 128K context, free self-hosted. Check the research license before commercial use.
Atlas with IBM Granite Code 20B (Ollama): More Capacity, Less Window in 2026
IBM Granite Code 20B (Ollama) is 12GB on disk and Free (self-hosted), but the 20b tag drops to 8K tokens (8,192) where the 8B instruct advertises 125K.
Atlas with MiniMax-M2.7-highspeed: The Fast Lane in 2026
MiniMax-M2.7-highspeed gives Atlas priority serving on MiniMax's newest 230B MoE at $0.60 per Mtok input and $2.40 per Mtok output, with a 204,800 token context.
Atlas with MiniMax-M3 in 2026: A Million Tokens for Thirty Cents
MiniMax-M3 gives Atlas a 1,000,000 token context at $0.30 per Mtok input and $1.20 per Mtok output, the cheapest large-context reasoning offer of 2026. Setup and limits.
Atlas with Mixtral 8x7B (Ollama): Sparse MoE Throughput in 2026
Run Atlas on Mixtral 8x7B (Ollama): a 26GB Apache 2.0 sparse mixture of experts with 32K context, free self-hosted. Memory math, throughput, and honest limits.
Atlas with Llama 3.1 8B (local via Ollama): The 4.9GB Baseline for 2026
Llama 3.1 8B (local via Ollama) in Atlas for 2026: a 4.9GB pull that fits 8GB of VRAM, Free (self-hosted), and honest limits on a general-purpose 8B model.
Atlas with Devstral Medium (2507): Agent-Trained Frontier Coding in 2026
Devstral Medium (2507) gives Atlas agent-first training at $0.40 / 1M input tokens and $2.00 / 1M output tokens, with 128,000 tokens in and out. Setup and tradeoffs.
Atlas with Mistral Nemo: The 128K Small Model Slot in 2026
Running Atlas on Mistral Nemo in 2026: 128,000 tokens of context in a 12B model at $0.15 / 1M input tokens, with the Tekken tokenizer that compresses code.
Atlas with Magistral 24B (Ollama): A Local Reasoning Model for the Plan Agent in 2026
Run Atlas on Magistral 24B (Ollama): Mistral's 14GB reasoning model with a 39K context, free self-hosted. Use it as the plan agent, then hand edits to a coder.
Atlas with Qwen3-Next 80B-A3B Instruct: Setup, Cost, and Tradeoffs in 2026
Run Atlas, the terminal-native AI coding agent, on Qwen3-Next 80B-A3B Instruct: 128K tokens (131,072) of context at $0.50 per Mtok input and $2.00 per Mtok output.
Atlas with Command R+ in 2026: Stronger Tool Use, 128K Context
Command R+ drives Atlas with a 128,000 token context and stronger multi step tool use, priced at $2.5 per Mtok input and $10 per Mtok output with a 4,000 token cap.
Atlas with Qwen3.5 Plus: A Million-Token Window for $0.40 per Mtok in 2026
Qwen3.5 Plus gives Atlas a 1M tokens (1,000,000) context window at $0.40 per Mtok input and $2.40 per Mtok output. What the million tokens buy, and what closed weights cost.
Atlas with Qwen3.6 27B: The 2026 Dense Checkpoint, Priced Honestly
Qwen3.6 27B is the dense reasoning model of Alibaba's April 2026 line: 256K tokens (262,144) of context at $0.60 per Mtok input and $3.60 per Mtok output, running inside Atlas.
Atlas with DeepSeek V4 Pro: 384K Output Tokens at $0.87 in 2026
DeepSeek V4 Pro in Atlas: the April 2026 flagship with a 1M token window, 384,000 max output tokens, and $0.435 / $0.87 per Mtok. Setup, data residency, tradeoffs.
Atlas with Amazon Nova Micro in 2026: The Cheapest Model on Bedrock
Amazon Nova Micro costs $0.035 per Mtok input, the lowest price in the Bedrock catalog, with a 128K token context. Use it as Atlas's small_model, never as the build loop.
Atlas with Command A Plus: Run a Frontier Model Inside Your Own VPC (2026)
Command A Plus drives Atlas at $2.50 / $10 per Mtok on a 128K window, and Cohere can deploy it in your own VPC or on-prem so code never leaves your network.
Atlas with Gemma 4 31B IT: Running Google's Open Weights Locally in 2026
Gemma 4 31B IT in Atlas: Google's April 2026 open-weights release with a 262,144 token window, free self-hosted via ollama pull gemma4:31b, or $0.99 / $1.49 on Cerebras.
Atlas with GPT-OSS 120B (local via Ollama): Offline Reasoning for Air-Gapped Work in 2026
GPT-OSS 120B is the strongest fully offline reasoning model for Atlas: a 65GB MXFP4 download, 131,072 token context, free self-hosted, or $0.15 / $0.60 per Mtok on Groq.
Atlas with Cloudflare Workers AI in 2026: The Cheapest Input Token in the Registry
Atlas with Cloudflare Workers AI in 2026: IBM Granite 4.0 H Micro at $0.017/$0.112 per Mtok, Kimi K2.7 Code at 262,144 tokens, and edge inference tradeoffs.
Atlas with NVIDIA NIM: Free-Tier Open Weights and the Nemotron Home Turf in 2026
Run Atlas on NVIDIA NIM: most endpoints listed at $0/$0 per Mtok, Nemotron 3 Ultra 550B at 1,000,000 tokens for $0.50/$2.50. Setup, rate limits, and catalog noise.
Atlas with Devstral Small 2 24B (local via Ollama): the Agent-First Local Model in 2026
Devstral Small 2 24B is Mistral's agent-first local model: a 14GB Ollama download, 128K context, free self-hosted, and it runs on a 16GB GPU. Atlas setup and tradeoffs.
Atlas with Ministral 8B: The Cheap Slot That Can Still Call Tools in 2026
Ministral 8B runs 128,000 tokens at $0.10 / 1M input tokens and $0.10 / 1M output tokens, symmetric. The Atlas small_model upgrade when subagents misfire on schemas.
Atlas with DeepSeek-R1 14B Distill (Ollama): The Local Plan Agent in 2026
DeepSeek-R1 14B Distill (Ollama) is 9.0GB, roughly 11GB to serve, with a 128K context. Drive the Atlas plan agent on a 12GB card in 2026. Free (self-hosted).
Atlas with Mistral NeMo 12B (Ollama): 128K Context on a 12GB Card in 2026
Run Atlas on Mistral NeMo 12B (Ollama): 7.1GB, a 128K practical context, free self-hosted. Why the Ollama tag says 1000K, and what limit.context to actually set.
Atlas with MiniMax-M2.7 in 2026: Agentic Reasoning at $0.30
MiniMax-M2.7 is MiniMax's March 2026 agentic 230B MoE. It runs Atlas at $0.30 per Mtok input and $1.20 per Mtok output with a 204,800 token context and 131,072 output.
Atlas with Qwen2.5-Coder 14B (Ollama): a real local build agent in 2026
Qwen2.5-Coder 14B (Ollama) in Atlas: 9.0GB of Q4_K_M weights, roughly 11GB to serve, 32K tokens (32,768) of context, Free (self-hosted), steady on tool chains.
Atlas with Claude Sonnet 4.5: The First 1M Token Claude in 2026
Claude Sonnet 4.5 gives Atlas a 1M token window at $3 per Mtok input, $15 per Mtok output. Setup, the 64K output ceiling, and when Sonnet 5 is the better pin.
Atlas with GLM-4.7 Flash: A Free 200K Context Model for the small_model Slot (2026)
GLM-4.7 Flash is free at $0 / $0 per Mtok with a 200,000 token context. Set it as Atlas's small_model so titles and summaries cost nothing at all.
Atlas with Llama 4 Maverick: 1M Context on Open Weights in 2026
Llama 4 Maverick gives Atlas a 1M token context on open weights at $0.24 / $0.97 per Mtok on Bedrock or $0.20 / $0.80 on DeepInfra. Setup, tradeoffs, and fit.
Atlas with Claude Haiku 4.5: The Cheap Slot in 2026
Claude Haiku 4.5 runs Atlas's small_model slot at $1 / $5 per Mtok with a 200K window. Titles, commit summaries, and cheap subagent fan-out, priced honestly for 2026.
Atlas with Qwen3.7 Max: Alibaba's May 2026 Frontier Tier at $2.50 / $7.50
Qwen3.7 Max in Atlas: Alibaba's May 2026 flagship, a 1M context model at $2.50 / $7.50 per Mtok, or $1.25 / $3.75 through Together. Setup with DASHSCOPE_API_KEY.
Atlas with Devstral 2: Mistral's Agent-First Coding Model in 2026
Devstral 2 in Atlas: Mistral's agent-trained coding model with a 262,144 token window at $0.40 / $2 per Mtok, loaded through @ai-sdk/mistral. Setup and tradeoffs.
Atlas with GLM-4.5-Flash: The Free Reasoning Slot in 2026
GLM-4.5-Flash is listed at $0.00 per Mtok input and output on Z.ai, with the full 128K tokens (131,072) context. Free background traffic for Atlas, rate limited.
Atlas with Gemini 2.0 Flash-Lite: The Cheapest Google Model in the Registry (2026)
Gemini 2.0 Flash-Lite in Atlas: $0.075 per Mtok input, $0.3 per Mtok output, a 1,048,576 token context, an 8,192 token output cap, and no reasoning mode.
Atlas with OpenAI o1-pro (2026): The $600 Per Mtok Question
OpenAI o1-pro is the most expensive model in the OpenAI registry at $150 per Mtok input and $600 per Mtok output. Here is what it does in Atlas and why o3 usually wins.
Atlas with NVIDIA Nemotron 3 Super 120B A12B in 2026
Nemotron 3 Super 120B A12B in Atlas, 2026: a reasoning MoE from $0.15/$0.65 per Mtok on Vercel AI Gateway, with 262,144 tokens of context on NVIDIA NIM.
Atlas with Nebius Token Factory in 2026: EU Infrastructure and the -fast Latency Lever
Atlas with Nebius Token Factory in 2026: EU-operated infrastructure, Qwen3.5-397B-A17B at $0.60/$3.60 per Mtok, and an 8,192 token output cap to plan around.
Atlas with Gemma 4 E4B (Ollama): the Default Gemma 4 Tag in 2026
Gemma 4 E4B (Ollama) is the :latest tag of Google's newest Gemma line: 9.6GB, 128K tokens (131,072), Free (self-hosted), with qat, mlx, mxfp8 and nvfp4 quants.
Atlas with Llama 3.1 70B (local via Ollama): A 43GB Private Agent in 2026
Llama 3.1 70B (local via Ollama) in Atlas for 2026: a 43GB pull for a 48GB GPU or 64GB Mac, Free (self-hosted), with 128,000 tokens of context on your own hardware.
Atlas with Mixtral 8x7B: Running the Original Sparse MoE in 2026
Mixtral 8x7B put sparse mixture-of-experts on the map in December 2023. In Atlas it gives 32,000 tokens at $0.70 / 1M input tokens. Setup, limits, and honest fit.
Atlas with Qwen3 Coder Flash: 1M Context at $0.30 per Mtok in 2026
Qwen3 Coder Flash in Atlas: 1M tokens (1,000,000) of context, 65,536 token output, $0.30 per Mtok input and $1.50 per Mtok output, tuned for fast edit loops.
Atlas with Qwen3 8B (local via Ollama): The Laptop Offline Setup for 2026
Qwen3 8B (local via Ollama) is the easiest offline Atlas setup in 2026: roughly 5.2 GB at Q4_K_M, Free (self-hosted), and a thinking model on a laptop GPU.
Atlas with GPT-4.1 nano (2026): A $0.10 Triage Model, Not a Coding Agent
GPT-4.1 nano runs in Atlas at $0.10 per Mtok input and $0.40 per Mtok output on a 1,047,576 token window. Use it for file triage and commit messages, not refactors.
Atlas with GPT-OSS 20B (local via Ollama): Local Reasoning on a 16GB Card in 2026
GPT-OSS 20B is OpenAI's open-weight reasoning model: 131,072 token context, runs on a single 16GB GPU, free self-hosted, or $0.075 / $0.30 per Mtok via Groq.
Atlas with DeepSeek-R1 (local via Ollama): Visible Reasoning at Every Size in 2026
DeepSeek-R1 runs locally from 1.5B to 671B under one Ollama tag, with a 128K to 164K context, free self-hosted, or $1.35 / $5.40 per Mtok via Bedrock. Atlas setup.
Atlas with GLM-4.6: 200K Open Weights at the 4.5 Price in 2026
GLM-4.6 runs Atlas on a 200K tokens (204,800) context at $0.60 per Mtok input and $2.20 per Mtok output, the same price GLM-4.5 charged on a 128K window.
Atlas with Qwen3.5 122B-A10B: The Middle MoE, Reviewed for 2026
Qwen3.5 122B-A10B brings 10B active parameters and 256K tokens (262,144) of context to Atlas at $0.40 per Mtok input and $3.20 per Mtok output. Setup, value, and where it loses.