Reference · 20 changes
LLM changelog
Every API model launch, price change, deprecation and retirement we track, newest first. Each entry links to the vendor's own notice; changes that affect what you pay get a short news write-up.
Last change on
Sep 2026
GPT-5.5 to leave ChatGPT, ChatGPT Work and Codex on October 14
Announced by OpenAI's @ChatGPT account. The OpenAI API is not affected: the deprecations page lists no shutdown for gpt-5.5.
OpenAI API pricingRead the storyindependentGizmochina, citing @ChatGPT
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available
Two audio-to-audio models for real-time voice applications through the Live API.
gpt-5.4-cyber deprecated, removal on October 1, 2026
Replacement: gpt-5.6-cyber.
DeepSeek V4.1 Flash released with new peak and off-peak prices
deepseek-flash costs $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak. V4 Flash and V4-Flash-Vision-Exp are retired; their names route to V4.1 Flash. A planned move of V4 Pro traffic to Flash from September 14 was later withdrawn, and V4 Pro stays at unchanged prices.
GPT-6 Astra released at $10 / $50 per 1M tokens
$20 / $75 above 272K input tokens. Batch and Flex at half price, Fast mode at double.
Gemini 3.8 Flash generally available at launch pricing
$0.75 / $3.75 per 1M tokens through December 31, 2026, then $1.50 / $7.50 from January 1, 2027.
Google Gemini API pricingRead the storyvendorGemini API release notes
Claude Fable 5.1 and Claude Mythos 5.1 launched
$10 / $50 per 1M tokens, the same as Fable 5. Cache reads cost $0.25, or 0.025x the input price instead of the usual 0.1x. Mythos 5.1 has limited availability.
Aug 2026
GPT-5.6 Sol cut to $4 / $20 on a promotion
20% lower input and 33% lower output. The promotional price runs at least through November 21, 2026.
Gemini 3.7 Flash generally available at an introductory price
$0.75 / $3.75 per 1M tokens through December 31, 2026. Google has not published the 2027 price.
DeepSeek V4 Pro generally available
The model version behind deepseek-v4-pro is DeepSeek-V4-Pro-0813.
Grok 4.6 released at $2 / $6 per 1M tokens
Cached input $0.50. Above 200K prompt tokens: $4 / $1 / $12. 500K context window.
Claude Sonnet 5 keeps $2 / $10; planned increase cancelled
The introductory price became the standard price. The increase to $3 / $15 scheduled for September 1, 2026 did not happen.
Anthropic API pricingRead the storyvendorClaude Platform release notes
Claude Opus 4.1 retired
Requests to claude-opus-4-1-20250805 now return an error.
Jul 2026
GPT-5.6 Luna 80% cheaper, Terra 20% cheaper; Fast mode replaces Priority
Luna now costs $0.20 / $1.20 and Terra $2 / $12. Fast mode costs twice the standard rate.
Claude Opus 5 launched at $5 / $25 per 1M tokens
1M-token context window, 128K max output, thinking on by default.
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available
Gemini 3.6 Flash costs $0.75 / $3.75 through December 31, 2026, then $1.50 / $7.50. Gemini 3.5 Flash-Lite costs $0.30 / $2.50.
GPT-5.6 family released: Sol, Terra and Luna
The gpt-5.6 alias routes to gpt-5.6-sol.
Grok 4.5 released at $2 / $6 per 1M tokens
Configurable reasoning effort: low, medium or high.
Jun 2026
Claude Sonnet 5 launched at introductory $2 / $10
1M-token context window and 128K max output tokens.
GPT-5 and o3 snapshots scheduled for removal on December 11, 2026
Includes gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07 and o3-2025-04-16. Replacements are the GPT-5.6 models.
How this log is kept
We watch vendor release notes, pricing pages and deprecation schedules, and use a news feed to catch changes vendors announce elsewhere first. An entry goes in only after it is confirmed on the vendor's own page, or, when the vendor announced it only on social media, in named press coverage that quotes the announcement. The date is the day the vendor published the change, not the day it takes effect.
Current prices for every model are on the LLM API pricing page.