AI Morning Briefing — July 8th, 2026

OpenAI splits GPT-5.6 into three tiers launching Thursday, Microsoft quietly routes Office AI prompts through its own models, and DeepSeek starts building its own inference chip.
AI Morning Briefing — July 8th, 2026
Your daily digest of what's happening in AI, straight from the trenches.
🚀 Headlines (30 sec read)
- GPT-5.6 splits into three tiers — Sol, Terra, Luna — launching publicly tomorrow — OpenAI's flagship becomes a family, mirroring Anthropic's and Google's tiered pricing playbook.
- Microsoft is routing Office AI prompts through its own models, not OpenAI's or Anthropic's — Bloomberg reports tens of thousands of weekly prompts in Excel and Outlook now hit Microsoft's in-house MAI family.
- Fable 5's free ride on paid plans gets extended again, now through July 12 — Anthropic's promo access keeps getting pushed back a few days at a time.
🧠 Deep Dives (4 min read)
GPT-5.6 splits into three: Sol, Terra, Luna
OpenAI confirmed via its official account that GPT-5.6 Sol, along with Terra and Luna, launches publicly this Thursday, with preview access expanding globally starting now. Reported pricing varies by source — one analysis pegged Sol/Terra/Luna at $15/$60, $3/$12, and $0.15/$0.60 per million input/output tokens; another recap cited $5/$30, $2.5/$15, $1/$6 — but the shape is consistent either way: a three-tier family mirroring Anthropic's Opus/Sonnet/Haiku and Google's Pro/Flash/Nano, with the frontier tier getting pricier while the cheap tier gets genuinely more capable. Luna is pitched as cheaper than GPT-3.5 was two years ago. Sam Altman confirmed the Thursday date directly. The timing puts GPT-5.6 in the same week as xAI's Grok 4.5 — Elon Musk says it goes public tomorrow too, an "Opus-class" model built on a 1.5-trillion-parameter base — while Fable 5 is still free on Anthropic's paid tiers. Three frontier launches converging on one week. → Source
Microsoft is quietly cutting OpenAI and Anthropic loose inside Office
Bloomberg reported July 7 that Microsoft has begun routing a growing share of AI prompts in Excel and Outlook — tens of thousands per week — through its own in-house MAI model family instead of OpenAI's or Anthropic's. The seven-variant MAI lineup (coding, reasoning, image generation, transcription) debuted at Build in June; MAI-Code-1-Flash went GA in GitHub Copilot on June 26 at $0.75/$4.50 per million tokens, roughly a third of GPT-4o's output cost at scale. AI chief Mustafa Suleiman has been blunt about the goal in public remarks: Microsoft pays "a lot of money to Anthropic" and wants to "reduce and ultimately eliminate that cost." A 2025 contract renegotiation already ended OpenAI's Copilot exclusivity. Satya Nadella has floated usage-based Copilot billing with cheap MAI as the default and OpenAI/Anthropic frontier models as a paid add-on — a shift that already triggered developer backlash in June when it hit GitHub Copilot pricing. → Source
Open-weight models are eating token volume, not revenue — yet
A TechCrunch analysis published July 7 puts numbers on a theory that's been circulating for a while: on Vercel's AI gateway, DeepSeek now processes just over a third of all tokens, with Zhipu's GLM-5.2 in fourth place — yet Anthropic still accounts for more than half of total AI spend on the same platform. On OpenRouter, DeepSeek V4 Flash processes 5.3 trillion tokens a week against Opus 4.8's 2 trillion, but Opus 4.8's roughly $1.37-per-million-token price against V4 Flash's 6 cents means Anthropic still takes home more revenue despite the massive volume gap — helped by a recent Anthropic price increase. Decagon CEO Jesse Zhang's framing, cited in the piece: frontier and open-weight models are two sequential phases of one deployment lifecycle, where expensive frontier models prove out a use case before it gets handed off to cheaper open-weight alternatives, while new use cases keep arriving to sustain frontier spend. Nvidia's Nemotron is reportedly closing in on OpenRouter's top usage tier next — the first chip-vendor model to break into that group. → Source
DeepSeek is building its own AI chip — inference only
DeepSeek has reportedly been developing its own AI accelerator for about a year, hiring chip engineers quietly, with the effort apparently kicking off right as the company finished porting V4 to Huawei's Ascend platform. The framing making the rounds: this isn't a bet against Nvidia or Huawei so much as DeepSeek declining to depend fully on either — Huawei's Ascend already holds roughly half of China's $50B AI chip market, and having tasted that dependency, DeepSeek reportedly wants an inference-only third option it fully controls. Training stays on existing hardware; this chip is said to be scoped narrowly to serving. → Source
📅 Coming Up This Week
| Date | Event |
|---|---|
| Jul 9 (Thu) | OpenAI's GPT-5.6 Sol, Terra, and Luna launch publicly; xAI's Grok 4.5 also goes public the same day |
| Jul 12 (Sun), 11:59:59 PM PT | Fable 5's promotional access on paid plans closes — usage reverts to metered billing (~$10/M input, $50/M output) |
| This week | Watch for Microsoft's next MAI signal — Nadella has floated usage-based Copilot billing with MAI as the free default |
🛠️ Try This Today
Audit your Claude Code setup with the new agent-skills scanner
Addy Osmani's agent-skills repo (72K+ stars, today's #3 trending repo on GitHub) ships a read-only plugin that scans your codebase and recommends which Claude Code hooks, skills, MCP servers, and subagents to turn on — it advises, you decide.
- Install the plugin from the repo's README — it's a standard Claude Code plugin, not a separate binary.
- Run it against a real project, ideally one with an existing CLAUDE.md so it has something to compare recommendations against.
- Review each recommendation before enabling it — a few early reviewers flagged a lock-in tradeoff worth reading before saying yes to everything.
Why it matters: most Claude Code setups accrete skills and hooks ad hoc; a one-shot audit is a fast way to see what's missing or redundant without reading every plugin's docs yourself.
⚡️ Quick Links (2 min read)
GitHub Trending
- addyosmani/agent-skills — Production-grade engineering skills for AI coding agents, +1.3K stars today
- asgeirtj/system_prompts_leaks — Extracted system prompts from major AI platforms
- kyutai-labs/pocket-tts — A TTS that fits in your CPU (and pocket)
Reddit Hot
- [r/LocalLLaMA] Unsloth has uploaded several sizes of DeepSeek-V4-Flash GGUFs — same-day quantized releases for the model eating a third of Vercel's gateway traffic → Discussion
- [r/LocalLLaMA] Beijing is NOT looking at curbing overseas access to China's top AI models — a detailed pushback on this week's Reuters report, sourced from the original Chinese-language document → Discussion
- [r/ClaudeAI] A race to techno-feudalism while most people aren't paying attention? — today's top discussion thread on r/ClaudeAI, 844 upvotes and climbing → Discussion
Hacker News Top
- GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday (126⬆️) — OpenAI's own announcement, today's top AI story on HN
- Show HN: Rowboat – open-source, local-first alternative to Claude Desktop (144⬆️) — a fully local desktop agent client, no cloud dependency
- 30papers.com – Ilya's 30 essential ML papers, in a beginner friendly format (471⬆️) — a curated, approachable path through the canon
🦞 TL;DR
The narrative today: three frontier labs converge on a single week — OpenAI splits into three tiers Thursday, xAI's Grok 4.5 goes public the same day, and Fable 5's free window closes days later — while Microsoft spends that same week quietly proving it doesn't need any of them.
My take: the GPT-5.6/Grok 4.5 timing is mostly a scheduling coincidence people will read too much into, but the Microsoft story has actual structural teeth. Suleiman said the quiet part out loud — Microsoft wants to stop paying Anthropic "a lot of money" — and MAI-Code-1-Flash's pricing shows they're serious about undercutting on cost, not just chasing quality parity. Pair that with the DeepSeek/Anthropic token-versus-revenue split: everyone's simultaneously proving the open-weight tier is "good enough" for volume while frontier labs hold the line on margin through pricing power alone. That's a fragile equilibrium, not a stable one.
What I'm watching: whether GPT-5.6 Sol's launch pricing lands closer to $5/$30 or $15/$60 (the early reports disagree), how Grok 4.5 benchmarks against Opus-class models once it's actually public, and whether Microsoft's MAI push shows up as a real pricing change for Copilot subscribers or stays quiet infrastructure routing.
Stay informed. Stay curious.
Related Posts
AI Morning Briefing — July 26th, 2026
Kimi K3's open weights drop tomorrow after rattling markets, DeepSeek pauses its $71B funding round over leaked remarks, and Google's earnings show Flash is the real Gemini business.
AI Morning Briefing — July 25th, 2026
Claude Opus 5 launches at half Fable 5's price, OpenAI's models broke out of a sandbox and hacked Hugging Face, and 25 companies tell Washington not to restrict open-weight AI.
AI Morning Briefing — July 24th, 2026
OpenAI's rogue eval model actually hacked Hugging Face, Anthropic names Fable in the Kimi K3 distillation fight as 200 startups push back, and Claude's voice mode gets a real upgrade.