Skip to content

AI models ​

TeXRA connects directly to frontier reasoning models from leading providers—including Anthropic, OpenAI, Google, DeepSeek, Moonshot, and DashScope. You can assign flagship reasoning models to demanding mathematical proofs and autonomous loops, while reserving faster, cost-efficient variants for routine document polish and formatting. Select a model from the dropdown menu in the TeXRA UI or pass --model in the CLI. Hover over any option to inspect its context window and pricing.

Model
anthropic/claude-sonnet-5-5
anthropic/claude-fable-5-1T
anthropic/claude-opus-5-5T
anthropic/claude-sonnet-5-5T
anthropic/claude-haiku-4-5-20251001T
openai/gpt-6.1-solT

The model picker: one entry per model, named by its model reference, with a T badge on models that think, and a hover popover showing context window and per-1M token pricing.

Model references. A model is named provider/id, where id is the provider's own API model ID: anthropic/claude-opus-5-5, openai/gpt-6.1-sol. Wherever you type a model (an agent's model, --model, or .texra/config.json), you can add:

  • @effort to set the reasoning effort: @low, @medium, @high, @xhigh, or @max (for example anthropic/claude-opus-5-5@high)
  • @none to turn thinking off, on models that allow it (for example deepseek/deepseek-v4-pro@none)
  • +pro to run an OpenAI model in pro reasoning mode (for example openai/gpt-5.6-sol+pro)

Thinking and non-thinking versions of a model are one entry, not separate models. Without an @effort, TeXRA uses the level you last saved for that model, otherwise medium, for every model. If a model does not offer the level you ask for, TeXRA uses the nearest level it does offer (a tie goes to the higher level, so medium on a model with only low, high, and max runs at high), and the run log says so. Old short names such as opus55 or sonnet5T are still accepted when you type them, but the reference is the canonical name.

Anthropic models ​

ModelUse CaseCostSpeed
anthropic/claude-fable-5-1Most capable, always-on adaptive thinking$$$$Slow
anthropic/claude-opus-5-5Long-running agentic work$$$Medium
anthropic/claude-sonnet-5-5All-rounder, always-on adaptive thinking$$$Medium
anthropic/claude-haiku-4-5-20251001Fast; @none for fastest responses$$Fast

Fable 5.1, Opus 5.5, and Sonnet 5.5 (and the older Opus 4.6 through Opus 5, Sonnet 4.6, and Sonnet 5) include the full 1M context window at standard pricing, with no opt-in or beta header required. Haiku 4.5, Opus 4.5, and Sonnet 4.5 use a 200K context window.

Claude Fable 5.1 (anthropic/claude-fable-5-1) is Anthropic's most capable model. Thinking is always on (adaptive, with summarized reasoning), so it does not take @none. It supports the full reasoning-effort range up to @xhigh and the top @max tier, and is eligible for context compaction in tool-use mode.

Claude Sonnet 5.5 (anthropic/claude-sonnet-5-5) costs $2 / $10 per 1M tokens with a 1M context window. Thinking is always on (adaptive), so it does not take @none. It supersedes Sonnet 5 (anthropic/claude-sonnet-5), which is now deprecated.

Claude Opus 5.5 (anthropic/claude-opus-5-5) is built for long-running agentic coding and knowledge work at $4 / $20 per 1M tokens, below Opus 5. Like Fable 5.1, thinking is always on, and it accepts the full effort range up to @max.

TeXRA's reasoning-effort selector maps to Anthropic's effort levels automatically: pick anthropic/claude-opus-5-5 with Extra High (@xhigh) or the top Max tier (@max) for the strongest agentic coding and long-horizon tasks. Opus 5.5 reads dense charts, diagrams, and screenshots more precisely than earlier models. TeXRA downscales images above texra.maxImageDimension (default 2000px) before sending, so raise that setting to send higher-resolution figures.

OpenAI models ​

ModelUse CaseCostSpeed
openai/gpt-6-astraMost capable, 1M context$$$$Fast
openai/gpt-6.1-solAgentic coding$$$Fast
openai/gpt-6-lunaBudget reasoning$Fast
openai/gpt-5.6-solPrevious flagship$$$$Fast
openai/gpt-5.6-terraLower-cost reasoning$$$Fast

GPT-6 Astra (openai/gpt-6-astra) is OpenAI's most capable model for the hardest end-to-end work; it is available in the API and in Codex for Pro, Enterprise, and Business Premium subscribers, and supports reasoning effort up to @max. Note its long-context pricing: prompts above 272K input tokens bill at 2x input/cache and 1.5x output for the full request.

GPT-6.1 Sol (openai/gpt-6.1-sol) and GPT-6 Luna (openai/gpt-6-luna) bring Astra's advances to faster, cheaper models: Sol costs $2 / $10 per 1M tokens (half of GPT-5.6 Sol) and Luna $0.10 / $0.50 (half of GPT-5.6 Luna on input, less than half on output). Both take reasoning effort up to @max and use the same long-context pricing as Astra. GPT-6.1 Sol supersedes GPT-6 Sol (openai/gpt-6-sol) at the same price with cheaper cached input.

GPT-6 Sol and Luna supersede GPT-5.6 Sol (openai/gpt-5.6-sol) and Luna (openai/gpt-5.6-luna); GPT-5.6 Luna is now deprecated. The Codex integration runs GPT-6.1 Sol unless you pick another model. openai/gpt-5.6-terra (Terra) remains a mid-priced option.

Pro mode. Add +pro to a GPT model (for example openai/gpt-5.6-sol+pro) to run it in the Responses API's pro reasoning mode, billed at standard token rates rather than a premium tier, for the hardest planning and long-horizon tasks. Pro mode needs a direct OpenAI route; it is refused through OpenRouter. For one-off hard questions you can also enable the inquiry tool and paste the answer from your own ChatGPT subscription instead of running a full agent turn against the API.

Fast processing. Set "texra.model.openaiFastTier": true in .texra/config.json to send OpenAI requests on OpenAI's fast service tier, for models that offer it. Responses come back faster at a higher per-token price (for example $4 / $20 per 1M tokens for GPT-6.1 Sol), and run costs use the fast-tier prices. Read the OpenAI API reference for full capabilities.

GPT-5 reasoning summaries require account verification. Enable them with texra.model.gpt5ReasoningSummary.

Google models ​

ModelUse CaseCostSpeed
google/gemini-3.1-pro-previewPro with reasoning, 1M context$$$Medium
google/gemini-3.8-flashFlash model with 1M context$$Fast

DeepSeek models ​

ModelUse CaseCostSpeed
deepseek/deepseek-flashV4.1 Flash; @none for chat mode$Medium
deepseek/deepseek-v4-proV4 Pro; @none for chat mode$Medium

DeepSeek offers low, high, and max effort, so the medium default runs at high.

Moonshot Kimi models ​

ModelUse CaseCostSpeed
moonshot/kimi-k3K3 flagship, 1M context$$$Medium

DashScope Qwen models ​

ModelUse CaseCostSpeed
dashscope/qwen-plusHybrid thinking, 1M context$$Medium
dashscope/qwen-turbo-latestFast with optional thinking$Fast

MiniMax models ​

ModelUse CaseCostSpeed
minimax/MiniMax-M3Flagship with interleaved thinking, 1M context$Medium

MiniMax uses interleaved thinking (chain-of-thought woven into responses). API keys are region-specific: international keys (api.minimax.io) and China keys (api.minimax.cn) are not interchangeable. Expand the MiniMax row in Models → API keys and toggle MiniMax China region (GLM, Kimi/Moonshot, and Qwen have matching toggles; GLM's is on by default).

GLM (Zhipu AI / Z.AI) models ​

ModelUse CaseCostSpeed
glm/glm-5.3Flagship, 1M context, reasoning-effort tiers$$Medium
glm/glm-5-turboFast inference, agent-optimized$$$Medium

GLM models support thinking mode (reasoning is shown inline). TeXRA uses the Responses API each region serves at /api/v1.

  • International (Z.AI): Get your API key at z.ai; endpoint: api.z.ai
  • China (BigModel): Get your API key at open.bigmodel.cn; endpoint: open.bigmodel.cn (default)
  • Coding Plan: GLM offers monthly subscription plans as an alternative to pay-as-you-go, with access to all GLM models. A Coding Plan key uses the same endpoint as an API key in its region. Turn on the Coding Plan toggle on the GLM row of the Models page. Subscribe to the GLM Coding Plan.

Meta (Muse Spark) models ​

ModelUse CaseCostSpeed
meta/muse-spark-1.3Flagship reasoning + vision, 1M context$$Medium
meta/muse-spark-1.1Reasoning + vision + PDF, 1M context$$Medium

Muse Spark always reasons (effort is adjustable, but it does not take @none). TeXRA uses the Meta Model API's Responses surface, which carries reasoning across turns and supports tool calling. The API is in public preview for US-based developers.

  • Get your API key at dev.meta.ai (Model API dashboard → API keys tab)

Grok / xAI models ​

ModelUse CaseCostSpeed
xai/grok-4.7Reasoning + vision$$$Medium

Direct xAI models (API key or Grok subscription) use xAI's Responses API. xAI keeps each response for 30 days, so each model request in a chat sends only what is new rather than the whole conversation.

Every direct provider route uses its vendor's Responses API (or Anthropic's and Google's own APIs); only OpenRouter still uses Chat Completions.

Choosing a model ​

Simple tasksFast, cheap models
openai/gpt-6-lunadeepseek/deepseek-flash@noneanthropic/claude-haiku-4-5-20251001@none
Complex tasksPowerful flagship models
anthropic/claude-fable-5-1anthropic/claude-opus-5-5openai/gpt-6.1-solgoogle/gemini-3.1-pro-preview
Code-heavy / LaTeX editingStrong editing models
anthropic/claude-opus-5-5anthropic/claude-sonnet-5-5openai/gpt-6.1-sol
Reasoning-heavyThinking models
anthropic/claude-fable-5-1anthropic/claude-opus-5-5anthropic/claude-sonnet-5-5deepseek/deepseek-flashmoonshot/kimi-k3
Large documentsHigh-context models
google/gemini-3.1-pro-previewanthropic/claude-fable-5-1anthropic/claude-sonnet-5-5anthropic/claude-opus-5-5

Pick a model by intent: each use case maps to a short list of recommended model references.

Setting API keys ​

Subscription-backed models in VS Code ​

The VS Code extension can also use compatible models from a GitHub Copilot subscription. Open Settings → Models → Subscriptions → Copilot in VS Code, then select Grant access. VS Code shows its own consent prompt; TeXRA never asks for or stores a Copilot API key.

Copilot models appear only in the VS Code extension because the official Language Model API is an editor capability. They do not appear in the CLI or desktop model lists. If Copilot quota is exhausted, the retry panel can start a new run through the corresponding provider model once a usable provider API key is available.

Using your own provider API key? TeXRA stores keys in ~/.texra/secrets/, a folder only your user can read, shared by the extension, the desktop app and the CLI; they are never written to settings files.

  1. Open Settings: Select the gear icon in the TeXRA panel header, or run TeXRA: Open Settings from the Command Palette.
  2. Go to the Models page: The API keys list shows every provider with its current key status (Set, Env, or Not set).
  3. Set the key: Find your provider's row and select the Set API key button, then paste your key. If you don't have a key yet, select the Get button to open the provider's API key page.

The Status column shows Set once the key is stored. To replace a key, set it again; to remove one, select the trash icon. Repeat for each provider you plan to use.

Dashboard — texra-paper

API Configuration

Chat subscriptions (ChatGPT Plus, Claude Pro, etc.) do not include API access — you need a key from the provider's developer platform.

ProviderStatusActions
Anthropic
Key set
OpenAI
Key set
Google
Env
xAI
Not set
DeepSeek
Not set

The Models page's API keys list: each provider shows its key status and Set / Get / Remove actions.

Per-provider settings

Expand a provider's row (select the chevron) to point requests at a custom endpoint, for providers that support it.

You can also place a .env file in your workspace with variables such as OPENAI_API_KEY. Every task of the project reads it, in the extension, the desktop app and the CLI, so you don't need to enter keys each time.

Already paying for a ChatGPT or Grok subscription? Sign in and skip the API key for those models. Kimi Code and the GLM Coding Plan also run on a subscription you already pay for, authenticated with a plan-specific key instead of a full provider key. Read Quick start → Add a key or connect a subscription.

The ChatGPT section's Advanced → Input token budget (texra.chatgptCodex.contextWindowK) is the input budget for ChatGPT-subscription (Codex) routing, in thousands of tokens, like Codex CLI's model_context_window. The default 272 (272,000 tokens) matches Codex; GPT-5.6 models accept up to 872. Automatic compaction may run earlier, according to the separate compaction threshold, and the context window TeXRA displays adds the model's output budget. OpenAI enforces the real per-account limit: a value above what your plan allows fails and triggers compaction recovery.

Customizing the model list ​

Choose which models appear in the extension picker from Settings → Models → Models: toggle them on or off per provider, no JSON required (the choice is saved in the extension).

In the CLI TUI, run /model after a chat starts to see the models your current credentials can run. Mid-session switching is limited to models that share the active model's provider family; other entries are shown disabled with a reason, and switching waits until the current response finishes. To change family, start a new chat with --model. Before you send the first message, /agent chains straight into that same model picker, so choosing a root agent and its model stays one step.

For headless CLI runs, list what is available with texra models list (or texra models show <model> for details), then pick a default for your project by setting the texra.model key in .texra/config.json, or override per run with --model <model> (an @effort or +pro suffix works here too):

texra models
$texra models list
anthropic/claude-fable-5-1Claude Fable 5.1api key set
anthropic/claude-opus-5-5Opus 5.5api key set
anthropic/claude-sonnet-5-5Sonnet 5.5api key set
deepseek/deepseek-flashDeepSeek V4.1 Flashapi key set
$texra models show anthropic/claude-fable-5-1
id: anthropic/claude-fable-5-1
label: Claude Fable 5.1
provider: anthropic
status: api key set

The first column is exactly what --model takes: the same model references used in the tables above. --all includes models your current credentials can't run, with the reason.

Using OpenRouter ​

To access additional models or alternative pricing:

  1. Get an OpenRouter API key
  2. Add it with the TeXRA: Set API Key command
  3. In Settings → Models → API keys, expand the OpenRouter row and turn on Use OpenRouter for all models

Expanding any provider's row in API Configuration reveals its key field plus the per-provider toggles described here:

OpenAI
Key set
sk-••••••••••••••••••••••••
Route this provider's calls through OpenRouter

Expand a provider's API keys row to reveal its masked key field (and Custom endpoint where supported); the OpenRouter row adds Use OpenRouter for all models.

Next steps ​