OpenAI and Anthropic announced steep price reductions for their mid‑tier large‑language models on Tuesday, a clear signal that U.S. AI labs are scrambling to keep pace with increasingly affordable Chinese alternatives. OpenAI’s newest GPT‑5.6 Luna model now costs $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6 respectively. Anthropic, meanwhile, rolled out Opus 5 at $5 per million input tokens and $25 per million output tokens, a 50 percent discount compared with its flagship Fable 5.
Both labs also adjusted their pricing roadmaps. Anthropic scrapped a planned increase for its Sonnet 5 model that was slated to take effect in September, while OpenAI left its top‑tier rates untouched, focusing the cuts on the middle of its product line.
How token pricing works
Customers are billed for the tokens that feed a model (input tokens) and the tokens the model generates (output tokens). The headline price per token, however, does not tell the whole story. More capable models can often finish a task with fewer tokens or fewer attempts, meaning a higher per‑token rate can translate into a lower overall cost for a given job.
Models also run at different "effort" settings, which adjust the computing power allocated to a request. Higher effort can improve performance but raises the cost per token. Artificial Analysis, a benchmark firm, found that Anthropic’s Opus 5 at a medium effort level delivered performance and cost per task comparable to Moonshot’s Kimi K3 at max effort. OpenAI’s GPT‑5.6 Luna at max effort performed on par with DeepSeek’s V4 Flash at the same effort level, but it cost just under twice as much per task.
Industry reaction
A source close to Anthropic said the lower pricing for Opus 5 reflects the company’s “family of models” strategy, not a direct attempt to match competitors. "We build a tiered lineup where each model serves a specific use case," the source explained.
Mantas Lukauskas, AI tech lead at web‑hosting provider Hostinger, noted that while prices for the most advanced models have stayed flat or risen, the new cuts represent the first real test of whether U.S. labs can protect the cost structure of their premium offerings. "The U.S. labs have cut the middle and are defending the top," Lukauskas said.
Both OpenAI and Anthropic declined to comment on the pricing changes. Analysts expect the price war to intensify as Chinese firms continue to roll out high‑performing models at lower price points, pushing U.S. providers to balance profitability with the need to retain developers who are increasingly cost‑sensitive.
For developers and enterprises that rely on large‑language models for everything from code generation to customer support, the new rates could translate into meaningful savings, especially for workloads that involve high token volumes. Yet the ultimate impact will depend on how quickly models can complete tasks more efficiently at lower effort settings, and whether the price cuts spur broader adoption or merely shift market share among existing players.
Questo articolo è stato scritto con l'assistenza dell'IA.
News Factory APP - notizie agentiche per potenziare il tuo SEO e AEO.