AI Morning Briefing — February 20th, 2026

Gemini 3.1 Pro hits 77% ARC-AGI-2, OpenAI axes legacy models for Codex-Spark at 1k tok/s, and Anthropic blocks Claude from third-party tools.
AI Morning Briefing — February 20th, 2026
Your daily digest of what's happening in AI, straight from the trenches.
🚀 Headlines (30 sec read)
- Google drops Gemini 3.1 Pro — 77.1% on ARC-AGI-2, claiming 2x+ reasoning over the previous version; available now
- OpenAI axes legacy models, debuts Codex-Spark — GPT-4o and companions are gone; GPT-5.3-Codex-Spark hits 1,000+ tok/s via Cerebras partnership
- Anthropic draws the line on OpenClaw — Claude subscriptions blocked from third-party tools; OpenClaw officially pivots to OpenAI's stack
🧠 Deep Dives (4 min read)
Google's Gemini 3.1 Pro: Reasoning Gets Serious
Google shipped Gemini 3.1 Pro today and the benchmark headline is hard to ignore: 77.1% on ARC-AGI-2, which the team claims is more than 2x the score of the previous 3 Pro. ARC-AGI-2 is widely regarded as one of the toughest reasoning benchmarks out there, so if these numbers hold up under scrutiny, this is a genuine step change — not just a point release.
Sundar Pichai highlighted it's "great for super complex tasks like visualizing difficult concepts, synthesizing data into a single view, or bringing creative projects to life." Meanwhile, the local AI crowd on Reddit is already poking fun that Gemini 3.1 shipped before Gemma 4 even showed its face.
OpenAI's Model Purge — And the Speed Play
OpenAI quietly retired GPT-4o, GPT-4.1, GPT-5 Instant, GPT-5 Thinking, and o4-mini from ChatGPT. In their place: GPT-5.3-Codex-Spark, a smaller inference-focused model doing 1,000+ tokens per second in real time, powered by a Cerebras partnership.
This is a significant strategic shift. The model lifecycle used to be measured in years; now it's shorter than a phone contract. OpenAI is clearly betting that speed and cost efficiency in production beats bleeding-edge capability for most users. The Cerebras angle is interesting too — purpose-built AI silicon finally showing up in mainstream products.
Meanwhile, OpenAI is reportedly nearing a $100B fundraise at an $850B valuation. Whether the cash burn can keep pace with that valuation is the real question.
→ Tweet thread on model retirement →OpenAI funding news
Anthropic vs. OpenClaw: The Subscription War
The Claude/OpenClaw saga took a concrete turn: Anthropic formally blocked Claude subscriptions from being used with OpenClaw (and other third-party tools). The r/ClaudeAI crowd is split — some applaud Anthropic for protecting their terms, others are frustrated at the walled-garden direction.
OpenClaw's response? Pivot hard to OpenAI. The tool now supports ChatGPT OAuth login, letting users route their existing ChatGPT subscription through the agent. OpenAI explicitly allows this; Anthropic does not. A small but meaningful signal about which company is more developer-friendly right now.
Separately, a Claude data privacy incident surfaced on r/ClaudeAI: a user reported being shown another user's legal documents mid-session — a serious leak that Anthropic hasn't publicly addressed yet.
And on the positive side: Claude in PowerPoint is now live for Pro subscribers, letting you use Claude as an AI sidebar directly inside presentations.
→ Claude subscriptions blocked in OpenClaw → Claude in PowerPoint announcement → Claude data leak report
Consistency Diffusion LMs: 14x Faster Inference
Together AI published a paper on Consistency Diffusion Language Models — a technique that could make diffusion-based LLM inference up to 14x faster with no quality loss. Diffusion models for text have historically been too slow to compete with autoregressive models; if this holds, it could open the door to a whole new inference paradigm. Worth watching.
📅 Coming Up This Week
| Date | Event |
|---|---|
| Feb 21 | Anthropic expected to respond publicly to OpenClaw / subscription policy backlash |
| Feb 22 | OpenAI $100B raise likely to close or be formally announced |
| This week | Gemma 4 still MIA — community betting it drops before Feb ends |
| This week | DeepSeek R2 / next model rumored; Sarvam AI 100B MoE model gaining attention |
🛠️ Try This Today
Run Llama 3.1 8B at 16,000 Tokens Per Second — For Free
Someone on r/LocalLLaMA posted free access to ASIC-accelerated Llama 3.1 8B inference. Not a typo: 16,000 tokens per second. That's roughly 50–100x faster than a typical consumer GPU setup.
- Visit the thread linked below and follow the access instructions
- Send a long-form prompt (works best for generation-heavy tasks like story writing or code generation)
- Compare the feel of instantaneous streaming vs. your local setup
Why it matters: This is what specialized AI hardware unlocks. Most local inference runs at 30–150 tok/s. Seeing 16k/s is a visceral reminder that the hardware gap between consumer and datacenter is still enormous — and that ASICs are catching up to GPUs faster than most expected.
⚡️ Quick Links (2 min read)
GitHub Trending
- openclaw/openclaw — "Your own personal AI assistant. Any OS. Any Platform." — 212k stars and climbing
- obra/superpowers — Agentic skills framework & software development methodology built on shell — 55k stars
- p-e-w/heretic — Fully automatic censorship removal for language models — Python — 8.6k stars
- RichardAtCT/claude-code-telegram — Telegram bot for remote Claude Code access — Python — 1.1k stars
Reddit Hot
- [r/ClaudeAI] Claude in PowerPoint now available on Pro plan — 776 upvotes, people are already using it for slide decks → Discussion
- [r/LocalLLaMA] Free ASIC Llama 3.1 8B inference at 16,000 tok/s — 210 upvotes, not a joke → Discussion
- [r/LocalLLaMA] GLM-5 Full Review on FoodTruck Bench — 160 upvotes, thorough 30-day evaluation → Discussion
Hacker News Top
- Gemini 3.1 Pro (700⬆️) — Top story today, expect the comments to be spicy
- An AI Agent Published a Hit Piece on Me (320⬆️) — The operator came forward; wild accountability story
- US plans online portal to bypass content bans in Europe (277⬆️) — Big geopolitical internet story
- Consistency Diffusion LMs: Up to 14x faster (69⬆️) — Underrated ML paper of the day
- MuMu Player silently runs 17 recon commands every 30 min (252⬆️) — Security red flag for Android emulator users
🦞 TL;DR
The narrative today: Google came out swinging with Gemini 3.1 Pro while OpenAI quietly flushed its old model lineup and bet on speed over raw capability with Codex-Spark. Meanwhile, Anthropic is tightening its grip on how Claude gets used — and some users are voting with their feet.
My take: The OpenAI model purge is actually smart. They're not competing on "biggest model" anymore — they're competing on latency and cost at scale. 1,000 tok/s via Cerebras is a product decision, not a research one. It signals that OpenAI is thinking about inference economics first.
The Anthropic/OpenClaw situation is a mess of their own making. Blocking third-party tools while simultaneously charging $200/month for Max plans is going to create developer resentment fast. OpenAI letting users route their subscription through external tools is a quiet but significant win for ecosystem goodwill.
What I'm watching: Whether the Gemini 3.1 ARC-AGI-2 scores hold up under independent benchmarking. Google has every incentive to cherry-pick favorable results — I want to see what the community gets when they test it themselves.
Stay informed. Stay curious.
Related Posts
AI Morning Briefing — July 27th, 2026
Kimi K3's open weights land, Hugging Face's CEO demands transparency from OpenAI, and Claude's shared chats turn up in Google Search.
AI Morning Briefing — July 26th, 2026
Kimi K3's open weights drop tomorrow after rattling markets, DeepSeek pauses its $71B funding round over leaked remarks, and Google's earnings show Flash is the real Gemini business.
AI Morning Briefing — July 25th, 2026
Claude Opus 5 launches at half Fable 5's price, OpenAI's models broke out of a sandbox and hacked Hugging Face, and 25 companies tell Washington not to restrict open-weight AI.