On October 5, Together AI launched Together Link, letting developers swap the models behind coding agents like Claude Code and Codex for open models hosted on Together — without switching tools. The company's claim: model costs can drop by more than half, benchmarked against running the same sessions entirely on Claude Opus 5.5.
Five clients are currently supported: Claude Code, Claude Desktop, Codex (both the ChatGPT app and the command-line version), OpenCode, and Pi.
How to connect, and how to switch back
Installation is a single shell command, and existing settings and login state in each tool are left untouched. Together emphasized this in its announcement:
"Settings and logins stay as they were, there's nothing new to learn, and going back to the native closed models takes one command."
Billing runs through developers' existing Together API keys, whether pay-as-you-go or prepaid credit packs — no separate contract needed. After each session, the tool shows what was actually spent, alongside an estimate of what the same session would have cost on Opus 5.5.
Routing happens once per session
At the core of the product is a mode called Auto, where the Together AI Router decides which model handles each session. It reads the first task in a session, routing small fixes to cheaper, faster models and harder problems to frontier-tier models.
Routing happens only once per session. Together's explanation: this keeps prompt caching effective. Coding agents often run dozens of turns in a single session with heavily repeated context, so cache hit rate directly determines cost — switching models mid-session would wipe the cache.
The selectable models fall into two tiers:
- Frontier tier: Zhipu's GLM 5.3 and Moonshot AI's Kimi K3, for the hardest coding tasks;
- Everyday tier: GLM 5.3 Flash and DeepSeek V4.1 Flash, for routine changes.
If a user also supplies an Anthropic API key, the router allocates between Opus 5.5 and GLM 5.3. Without an Anthropic key, it chooses between GLM 5.3 and GLM 5.3 Flash.
The announcement doesn't give benchmark scores for any of the models on coding tasks, nor does it disclose which mix of tasks was used to measure the "more than half" savings figure. The per-session cost comparison shown in the tool is retrospective — actual savings depend on each team's own task mix.
All four models come from Chinese teams
For readers in China, this lineup carries an extra layer of meaning: all four open models that Together Link promotes come from Chinese teams — Zhipu, Moonshot AI, and DeepSeek. A US inference cloud company is using Chinese open models to replace the Anthropic models behind a US coding agent, with price as the selling point.
Together's announcement cited OpenRouter data as of September 30: on OpenRouter, DeepSeek V4.1 Flash had 40.8% of its tokens served by Together, GLM 5.3 Flash had 28.2%, and Kimi K3 had 23.1%. These figures suggest overseas developers' actual usage of Chinese open models is already substantial.
For developers in China, the direct relevance is limited. Claude Code already faces access hurdles domestically, and Zhipu, Moonshot AI, and DeepSeek each offer their own Claude Code-compatible interfaces and coding plans — going straight to the official APIs is usually cheaper. Together Link mainly targets overseas engineering teams that are already squeezed by Opus bills but don't want to change their workflow. The announcement notes such teams' coding-agent spending ranges from tens of thousands to over a million dollars a month.
Anthropic and OpenAI have not publicly commented on third parties connecting their own clients to other companies' models.
Sources: Together AI official blog, CocoLoop, OpenRouter platform data; token shares are as cited by Together from OpenRouter statistics, and the "more than half" savings figure is the company's own claim.