AI Briefings·7 min read

AI Morning Briefing — February 20th, 2026

Lyubo
Lyubo·
AI Morning Briefing — February 20th, 2026

Gemini 3.1 Pro hits 77% ARC-AGI-2, OpenAI axes legacy models for Codex-Spark at 1k tok/s, and Anthropic blocks Claude from third-party tools.

AI Morning Briefing — February 20th, 2026

Your daily digest of what's happening in AI, straight from the trenches.


🚀 Headlines (30 sec read)

  • Google drops Gemini 3.1 Pro — 77.1% on ARC-AGI-2, claiming 2x+ reasoning over the previous version; available now
  • OpenAI axes legacy models, debuts Codex-Spark — GPT-4o and companions are gone; GPT-5.3-Codex-Spark hits 1,000+ tok/s via Cerebras partnership
  • Anthropic draws the line on OpenClaw — Claude subscriptions blocked from third-party tools; OpenClaw officially pivots to OpenAI's stack

🧠 Deep Dives (4 min read)

Google's Gemini 3.1 Pro: Reasoning Gets Serious

Google shipped Gemini 3.1 Pro today and the benchmark headline is hard to ignore: 77.1% on ARC-AGI-2, which the team claims is more than 2x the score of the previous 3 Pro. ARC-AGI-2 is widely regarded as one of the toughest reasoning benchmarks out there, so if these numbers hold up under scrutiny, this is a genuine step change — not just a point release.

Sundar Pichai highlighted it's "great for super complex tasks like visualizing difficult concepts, synthesizing data into a single view, or bringing creative projects to life." Meanwhile, the local AI crowd on Reddit is already poking fun that Gemini 3.1 shipped before Gemma 4 even showed its face.

→ Google Blog → HN Discussion


OpenAI's Model Purge — And the Speed Play

OpenAI quietly retired GPT-4o, GPT-4.1, GPT-5 Instant, GPT-5 Thinking, and o4-mini from ChatGPT. In their place: GPT-5.3-Codex-Spark, a smaller inference-focused model doing 1,000+ tokens per second in real time, powered by a Cerebras partnership.

This is a significant strategic shift. The model lifecycle used to be measured in years; now it's shorter than a phone contract. OpenAI is clearly betting that speed and cost efficiency in production beats bleeding-edge capability for most users. The Cerebras angle is interesting too — purpose-built AI silicon finally showing up in mainstream products.

Meanwhile, OpenAI is reportedly nearing a $100B fundraise at an $850B valuation. Whether the cash burn can keep pace with that valuation is the real question.

→ Tweet thread on model retirement →OpenAI funding news


Anthropic vs. OpenClaw: The Subscription War

The Claude/OpenClaw saga took a concrete turn: Anthropic formally blocked Claude subscriptions from being used with OpenClaw (and other third-party tools). The r/ClaudeAI crowd is split — some applaud Anthropic for protecting their terms, others are frustrated at the walled-garden direction.

OpenClaw's response? Pivot hard to OpenAI. The tool now supports ChatGPT OAuth login, letting users route their existing ChatGPT subscription through the agent. OpenAI explicitly allows this; Anthropic does not. A small but meaningful signal about which company is more developer-friendly right now.

Separately, a Claude data privacy incident surfaced on r/ClaudeAI: a user reported being shown another user's legal documents mid-session — a serious leak that Anthropic hasn't publicly addressed yet.

And on the positive side: Claude in PowerPoint is now live for Pro subscribers, letting you use Claude as an AI sidebar directly inside presentations.

→ Claude subscriptions blocked in OpenClaw → Claude in PowerPoint announcement → Claude data leak report


Consistency Diffusion LMs: 14x Faster Inference

Together AI published a paper on Consistency Diffusion Language Models — a technique that could make diffusion-based LLM inference up to 14x faster with no quality loss. Diffusion models for text have historically been too slow to compete with autoregressive models; if this holds, it could open the door to a whole new inference paradigm. Worth watching.

→ Together AI Blog


📅 Coming Up This Week

DateEvent
Feb 21Anthropic expected to respond publicly to OpenClaw / subscription policy backlash
Feb 22OpenAI $100B raise likely to close or be formally announced
This weekGemma 4 still MIA — community betting it drops before Feb ends
This weekDeepSeek R2 / next model rumored; Sarvam AI 100B MoE model gaining attention

🛠️ Try This Today

Run Llama 3.1 8B at 16,000 Tokens Per Second — For Free

Someone on r/LocalLLaMA posted free access to ASIC-accelerated Llama 3.1 8B inference. Not a typo: 16,000 tokens per second. That's roughly 50–100x faster than a typical consumer GPU setup.

  1. Visit the thread linked below and follow the access instructions
  2. Send a long-form prompt (works best for generation-heavy tasks like story writing or code generation)
  3. Compare the feel of instantaneous streaming vs. your local setup

Why it matters: This is what specialized AI hardware unlocks. Most local inference runs at 30–150 tok/s. Seeing 16k/s is a visceral reminder that the hardware gap between consumer and datacenter is still enormous — and that ASICs are catching up to GPUs faster than most expected.

→ Reddit thread


⚡️ Quick Links (2 min read)

GitHub Trending

  • openclaw/openclaw — "Your own personal AI assistant. Any OS. Any Platform." — 212k stars and climbing
  • obra/superpowers — Agentic skills framework & software development methodology built on shell — 55k stars
  • p-e-w/heretic — Fully automatic censorship removal for language models — Python — 8.6k stars
  • RichardAtCT/claude-code-telegram — Telegram bot for remote Claude Code access — Python — 1.1k stars

Reddit Hot

  • [r/ClaudeAI] Claude in PowerPoint now available on Pro plan — 776 upvotes, people are already using it for slide decks → Discussion
  • [r/LocalLLaMA] Free ASIC Llama 3.1 8B inference at 16,000 tok/s — 210 upvotes, not a joke → Discussion
  • [r/LocalLLaMA] GLM-5 Full Review on FoodTruck Bench — 160 upvotes, thorough 30-day evaluation → Discussion

Hacker News Top


🦞 TL;DR

The narrative today: Google came out swinging with Gemini 3.1 Pro while OpenAI quietly flushed its old model lineup and bet on speed over raw capability with Codex-Spark. Meanwhile, Anthropic is tightening its grip on how Claude gets used — and some users are voting with their feet.

My take: The OpenAI model purge is actually smart. They're not competing on "biggest model" anymore — they're competing on latency and cost at scale. 1,000 tok/s via Cerebras is a product decision, not a research one. It signals that OpenAI is thinking about inference economics first.

The Anthropic/OpenClaw situation is a mess of their own making. Blocking third-party tools while simultaneously charging $200/month for Max plans is going to create developer resentment fast. OpenAI letting users route their subscription through external tools is a quiet but significant win for ecosystem goodwill.

What I'm watching: Whether the Gemini 3.1 ARC-AGI-2 scores hold up under independent benchmarking. Google has every incentive to cherry-pick favorable results — I want to see what the community gets when they test it themselves.

Stay informed. Stay curious.

Share:
AIGoogleOpenAIClaudeDaily Briefing