AI Briefings·8 min read

AI Morning Briefing — July 8th, 2026

Lyubo
Lyubo·
AI Morning Briefing — July 8th, 2026

OpenAI splits GPT-5.6 into three tiers launching Thursday, Microsoft quietly routes Office AI prompts through its own models, and DeepSeek starts building its own inference chip.

AI Morning Briefing — July 8th, 2026

Your daily digest of what's happening in AI, straight from the trenches.


🚀 Headlines (30 sec read)

  • GPT-5.6 splits into three tiers — Sol, Terra, Luna — launching publicly tomorrow — OpenAI's flagship becomes a family, mirroring Anthropic's and Google's tiered pricing playbook.
  • Microsoft is routing Office AI prompts through its own models, not OpenAI's or Anthropic's — Bloomberg reports tens of thousands of weekly prompts in Excel and Outlook now hit Microsoft's in-house MAI family.
  • Fable 5's free ride on paid plans gets extended again, now through July 12 — Anthropic's promo access keeps getting pushed back a few days at a time.

🧠 Deep Dives (4 min read)

GPT-5.6 splits into three: Sol, Terra, Luna

OpenAI confirmed via its official account that GPT-5.6 Sol, along with Terra and Luna, launches publicly this Thursday, with preview access expanding globally starting now. Reported pricing varies by source — one analysis pegged Sol/Terra/Luna at $15/$60, $3/$12, and $0.15/$0.60 per million input/output tokens; another recap cited $5/$30, $2.5/$15, $1/$6 — but the shape is consistent either way: a three-tier family mirroring Anthropic's Opus/Sonnet/Haiku and Google's Pro/Flash/Nano, with the frontier tier getting pricier while the cheap tier gets genuinely more capable. Luna is pitched as cheaper than GPT-3.5 was two years ago. Sam Altman confirmed the Thursday date directly. The timing puts GPT-5.6 in the same week as xAI's Grok 4.5 — Elon Musk says it goes public tomorrow too, an "Opus-class" model built on a 1.5-trillion-parameter base — while Fable 5 is still free on Anthropic's paid tiers. Three frontier launches converging on one week. → Source

Microsoft is quietly cutting OpenAI and Anthropic loose inside Office

Bloomberg reported July 7 that Microsoft has begun routing a growing share of AI prompts in Excel and Outlook — tens of thousands per week — through its own in-house MAI model family instead of OpenAI's or Anthropic's. The seven-variant MAI lineup (coding, reasoning, image generation, transcription) debuted at Build in June; MAI-Code-1-Flash went GA in GitHub Copilot on June 26 at $0.75/$4.50 per million tokens, roughly a third of GPT-4o's output cost at scale. AI chief Mustafa Suleiman has been blunt about the goal in public remarks: Microsoft pays "a lot of money to Anthropic" and wants to "reduce and ultimately eliminate that cost." A 2025 contract renegotiation already ended OpenAI's Copilot exclusivity. Satya Nadella has floated usage-based Copilot billing with cheap MAI as the default and OpenAI/Anthropic frontier models as a paid add-on — a shift that already triggered developer backlash in June when it hit GitHub Copilot pricing. → Source

Open-weight models are eating token volume, not revenue — yet

A TechCrunch analysis published July 7 puts numbers on a theory that's been circulating for a while: on Vercel's AI gateway, DeepSeek now processes just over a third of all tokens, with Zhipu's GLM-5.2 in fourth place — yet Anthropic still accounts for more than half of total AI spend on the same platform. On OpenRouter, DeepSeek V4 Flash processes 5.3 trillion tokens a week against Opus 4.8's 2 trillion, but Opus 4.8's roughly $1.37-per-million-token price against V4 Flash's 6 cents means Anthropic still takes home more revenue despite the massive volume gap — helped by a recent Anthropic price increase. Decagon CEO Jesse Zhang's framing, cited in the piece: frontier and open-weight models are two sequential phases of one deployment lifecycle, where expensive frontier models prove out a use case before it gets handed off to cheaper open-weight alternatives, while new use cases keep arriving to sustain frontier spend. Nvidia's Nemotron is reportedly closing in on OpenRouter's top usage tier next — the first chip-vendor model to break into that group. → Source

DeepSeek is building its own AI chip — inference only

DeepSeek has reportedly been developing its own AI accelerator for about a year, hiring chip engineers quietly, with the effort apparently kicking off right as the company finished porting V4 to Huawei's Ascend platform. The framing making the rounds: this isn't a bet against Nvidia or Huawei so much as DeepSeek declining to depend fully on either — Huawei's Ascend already holds roughly half of China's $50B AI chip market, and having tasted that dependency, DeepSeek reportedly wants an inference-only third option it fully controls. Training stays on existing hardware; this chip is said to be scoped narrowly to serving. → Source


📅 Coming Up This Week

DateEvent
Jul 9 (Thu)OpenAI's GPT-5.6 Sol, Terra, and Luna launch publicly; xAI's Grok 4.5 also goes public the same day
Jul 12 (Sun), 11:59:59 PM PTFable 5's promotional access on paid plans closes — usage reverts to metered billing (~$10/M input, $50/M output)
This weekWatch for Microsoft's next MAI signal — Nadella has floated usage-based Copilot billing with MAI as the free default

🛠️ Try This Today

Audit your Claude Code setup with the new agent-skills scanner

Addy Osmani's agent-skills repo (72K+ stars, today's #3 trending repo on GitHub) ships a read-only plugin that scans your codebase and recommends which Claude Code hooks, skills, MCP servers, and subagents to turn on — it advises, you decide.

  1. Install the plugin from the repo's README — it's a standard Claude Code plugin, not a separate binary.
  2. Run it against a real project, ideally one with an existing CLAUDE.md so it has something to compare recommendations against.
  3. Review each recommendation before enabling it — a few early reviewers flagged a lock-in tradeoff worth reading before saying yes to everything.

Why it matters: most Claude Code setups accrete skills and hooks ad hoc; a one-shot audit is a fast way to see what's missing or redundant without reading every plugin's docs yourself.


⚡️ Quick Links (2 min read)

GitHub Trending

Reddit Hot

  • [r/LocalLLaMA] Unsloth has uploaded several sizes of DeepSeek-V4-Flash GGUFs — same-day quantized releases for the model eating a third of Vercel's gateway traffic → Discussion
  • [r/LocalLLaMA] Beijing is NOT looking at curbing overseas access to China's top AI models — a detailed pushback on this week's Reuters report, sourced from the original Chinese-language document → Discussion
  • [r/ClaudeAI] A race to techno-feudalism while most people aren't paying attention? — today's top discussion thread on r/ClaudeAI, 844 upvotes and climbing → Discussion

Hacker News Top


🦞 TL;DR

The narrative today: three frontier labs converge on a single week — OpenAI splits into three tiers Thursday, xAI's Grok 4.5 goes public the same day, and Fable 5's free window closes days later — while Microsoft spends that same week quietly proving it doesn't need any of them.

My take: the GPT-5.6/Grok 4.5 timing is mostly a scheduling coincidence people will read too much into, but the Microsoft story has actual structural teeth. Suleiman said the quiet part out loud — Microsoft wants to stop paying Anthropic "a lot of money" — and MAI-Code-1-Flash's pricing shows they're serious about undercutting on cost, not just chasing quality parity. Pair that with the DeepSeek/Anthropic token-versus-revenue split: everyone's simultaneously proving the open-weight tier is "good enough" for volume while frontier labs hold the line on margin through pricing power alone. That's a fragile equilibrium, not a stable one.

What I'm watching: whether GPT-5.6 Sol's launch pricing lands closer to $5/$30 or $15/$60 (the early reports disagree), how Grok 4.5 benchmarks against Opus-class models once it's actually public, and whether Microsoft's MAI push shows up as a real pricing change for Copilot subscribers or stays quiet infrastructure routing.

Stay informed. Stay curious.

Share:
AIOpenAIClaudeDaily Briefing