Reference · 20 changes

LLM changelog

Every API model launch, price change, deprecation and retirement we track, newest first. Each entry links to the vendor's own notice; changes that affect what you pay get a short news write-up.

Last change on

Sep 2026

  1. Retirement

    GPT-5.5 to leave ChatGPT, ChatGPT Work and Codex on October 14

    Announced by OpenAI's @ChatGPT account. The OpenAI API is not affected: the deprecations page lists no shutdown for gpt-5.5.

  2. New model

    Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available

    Two audio-to-audio models for real-time voice applications through the Live API.

  3. Deprecation

    gpt-5.4-cyber deprecated, removal on October 1, 2026

    Replacement: gpt-5.6-cyber.

  4. New model

    DeepSeek V4.1 Flash released with new peak and off-peak prices

    deepseek-flash costs $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak. V4 Flash and V4-Flash-Vision-Exp are retired; their names route to V4.1 Flash. A planned move of V4 Pro traffic to Flash from September 14 was later withdrawn, and V4 Pro stays at unchanged prices.

  5. New model

    GPT-6 Astra released at $10 / $50 per 1M tokens

    $20 / $75 above 272K input tokens. Batch and Flex at half price, Fast mode at double.

  6. New model

    Gemini 3.8 Flash generally available at launch pricing

    $0.75 / $3.75 per 1M tokens through December 31, 2026, then $1.50 / $7.50 from January 1, 2027.

  7. New model

    Claude Fable 5.1 and Claude Mythos 5.1 launched

    $10 / $50 per 1M tokens, the same as Fable 5. Cache reads cost $0.25, or 0.025x the input price instead of the usual 0.1x. Mythos 5.1 has limited availability.

Aug 2026

  1. Price cut

    GPT-5.6 Sol cut to $4 / $20 on a promotion

    20% lower input and 33% lower output. The promotional price runs at least through November 21, 2026.

  2. New model

    Gemini 3.7 Flash generally available at an introductory price

    $0.75 / $3.75 per 1M tokens through December 31, 2026. Google has not published the 2027 price.

  3. New model

    DeepSeek V4 Pro generally available

    The model version behind deepseek-v4-pro is DeepSeek-V4-Pro-0813.

  4. New model

    Grok 4.6 released at $2 / $6 per 1M tokens

    Cached input $0.50. Above 200K prompt tokens: $4 / $1 / $12. 500K context window.

  5. Price change

    Claude Sonnet 5 keeps $2 / $10; planned increase cancelled

    The introductory price became the standard price. The increase to $3 / $15 scheduled for September 1, 2026 did not happen.

  6. Retirement

    Claude Opus 4.1 retired

    Requests to claude-opus-4-1-20250805 now return an error.

Jul 2026

  1. Price cut

    GPT-5.6 Luna 80% cheaper, Terra 20% cheaper; Fast mode replaces Priority

    Luna now costs $0.20 / $1.20 and Terra $2 / $12. Fast mode costs twice the standard rate.

  2. New model

    Claude Opus 5 launched at $5 / $25 per 1M tokens

    1M-token context window, 128K max output, thinking on by default.

  3. New model

    Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available

    Gemini 3.6 Flash costs $0.75 / $3.75 through December 31, 2026, then $1.50 / $7.50. Gemini 3.5 Flash-Lite costs $0.30 / $2.50.

  4. New model

    GPT-5.6 family released: Sol, Terra and Luna

    The gpt-5.6 alias routes to gpt-5.6-sol.

  5. New model

    Grok 4.5 released at $2 / $6 per 1M tokens

    Configurable reasoning effort: low, medium or high.

Jun 2026

  1. New model

    Claude Sonnet 5 launched at introductory $2 / $10

    1M-token context window and 128K max output tokens.

  2. Deprecation

    GPT-5 and o3 snapshots scheduled for removal on December 11, 2026

    Includes gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07 and o3-2025-04-16. Replacements are the GPT-5.6 models.

How this log is kept

We watch vendor release notes, pricing pages and deprecation schedules, and use a news feed to catch changes vendors announce elsewhere first. An entry goes in only after it is confirmed on the vendor's own page, or, when the vendor announced it only on social media, in named press coverage that quotes the announcement. The date is the day the vendor published the change, not the day it takes effect.

Current prices for every model are on the LLM API pricing page.