AI Briefings·5 min read

AI Morning Briefing — May 8th, 2026

Lyubo
Lyubo·
AI Morning Briefing — May 8th, 2026

Anthropic rents Colossus 1 (220K GPUs), OpenAI launches real-time voice translation, and DeepSeek V4 hits a $45B valuation

AI Morning Briefing — May 8th, 2026

Your daily digest of what's happening in AI, straight from the trenches.


🚀 Headlines (30 sec read)

  • Anthropic rents Colossus 1 — 220,000 NVIDIA GPUs from SpaceX/xAI; Claude Code limits doubled for Pro/Max/Team
  • OpenAI launches GPT-Realtime-2 — GPT-5-class voice model with 128K context and live 70+ language translation
  • DeepSeek V4 hits $45B valuation — China's state fund leads debut fundraise; V4 tops charts on day one

🧠 Deep Dives (4 min read)

Anthropic Plugs Into Space: The Colossus Compute Deal

Anthropic has signed a landmark compute agreement with SpaceX's xAI, gaining full access to Colossus 1 — a 220,000+ GPU supercomputer drawing 300 megawatts. The immediate user impact: Claude Code's 5-hour limit is now doubled across Pro, Max, and Team plans, peak-hour throttling is eliminated for Pro and Max subscribers, and Opus API limits have jumped significantly.

The quietly stunning part is what Anthropic said next: it is already eyeing "multi-gigawatt compute in space." Six months ago xAI and Anthropic were pure rivals. Today one is renting the other's supercomputer because the real bottleneck has shifted from chips to electricity. The race just went orbital. → Source

OpenAI's Voice AI Goes Real-Time and Multilingual

OpenAI announced three new real-time audio models anchored by GPT-Realtime-2 and GPT-Realtime-Translate. GPT-Realtime-2 brings GPT-5-class reasoning to voice, expands context from 32K to 128K tokens, and adds parallel tool-calling with spoken status narration ("I'm checking your calendar…"). It also introduces tunable reasoning levels (minimal → xhigh) and graceful failure responses instead of awkward silences.

GPT-Realtime-Translate goes further: live speech-to-speech translation across 70+ input languages into 13 output languages, synced to the speaker's pace. No pause-and-replay. No upload-and-wait. This is the language barrier thinning in real time — and it will quietly hollow out a large chunk of the language services industry. → Source

DeepSeek V4 Tops Charts; $45B Valuation Signals Sovereign AI

DeepSeek V4 launched and immediately topped the model leaderboards. More consequentially, China's state semiconductor "Big Fund" is reportedly negotiating to lead DeepSeek's debut fundraising round at a $45 billion valuation — a figure that would rank it alongside the world's top AI labs. The shift from venture capital to sovereign capital in pricing frontier AI marks a structural change: these models are now considered strategic national infrastructure, not just tech products.

Dario Amodei this week separately revealed Anthropic has grown 80x. The gap between the top three labs (OpenAI, Anthropic, DeepSeek) and everyone else is widening fast. → Source


📅 Coming Up This Week

DateEvent
May 13 (Wed)Nous Research AMA on r/LocalLLaMA — 8AM–11AM PST, covering their open-source Hermes Agent
This weekDeepSeek fundraising round expected to close — watch for official valuation announcement
NowOpenAI GPT-Realtime-2 API live for developers — early access to real-time voice + translation

🛠️ Try This Today

Enable Multi-Token Prediction in LLaMA.cpp for 40% Faster Local Inference

A community tutorial on r/LocalLLaMA shows how to unlock Multi-Token Prediction (MTP) in the latest LLaMA.cpp — getting ~40% faster generation with Gemma 4 models at no quality cost:

  1. Pull the latest LLaMA.cpp (MTP support merged this week)
  2. Add --mtp-n-draft 4 to your inference command
  3. Benchmark with --n-predict 200 before and after to confirm speedup

Why it matters: MTP is a form of speculative decoding — the model predicts multiple tokens ahead and verifies in parallel. It cuts wall-clock latency dramatically for longer generations, which makes local AI meaningfully more practical for conversational and coding workloads. → Original post


⚡️ Quick Links (2 min read)

GitHub Trending

Reddit Hot

  • [r/LocalLLaMA] Multi-Token Prediction in LLaMA.cpp speeds up Gemma 4 by 40% — Tutorial + benchmark thread worth bookmarking → Discussion
  • [r/LocalLLaMA] WARNING: Open-OSS/privacy-filter MALWARE on Hugging Face — Infostealer disguised as an OpenAI privacy filter; Windows users must avoid → Discussion
  • [r/LocalLLaMA] Skymizer HTX301: 384GB VRAM PCIe card at 240W — New Taiwanese inference hardware for running very large models locally → Discussion

Hacker News Top


🦞 TL;DR

The narrative today: Compute is the new oil, and Anthropic just secured a gusher — 220,000 GPUs from SpaceX while simultaneously racing DeepSeek and OpenAI on capability.

My take: The Colossus deal is being underreported. Anthropic going straight to a competitor's supercomputer tells you the electricity bottleneck is more real than the chip shortage ever was. OpenAI's voice translation is genuinely impressive and will quietly eat the language services industry over the next two years. And the DeepSeek $45B valuation is less about the number and more about what it signals: sovereign funds now price frontier AI as national infrastructure. The age of scrappy AI startups raising Series A is ending.

What I'm watching: Whether Anthropic's "multi-gigawatt compute in space" comment is a roadmap item or marketing hyperbole. If it's real, AI infrastructure timescales just changed fundamentally.

Stay informed. Stay curious.

Share:
AIOpenAIAnthropicDeepSeekDaily Briefing