AI Briefings·8 min read

AI Morning Briefing — May 26th, 2026

Lyubo
Lyubo·
AI Morning Briefing — May 26th, 2026

DeepSeek cuts prices 75% permanently, Anthropic's 'Dreaming' agents learn from failure, mystery Mythos 1 model surfaces, and the famous METR benchmark gets torn apart.

AI Morning Briefing — May 26th, 2026

Your daily digest of what's happening in AI, straight from the trenches.


🚀 Headlines (30 sec read)

  • DeepSeek permanently slashes V4-Pro prices by 75% — Not a promo. The commoditization floor keeps dropping, and it's accelerating.
  • Anthropic debuts "Dreaming" — Claude agents that learn from past failures — Revealed at Code with Claude 2026; Harvey saw 6x task completion gains.
  • Mystery "Mythos 1" model spotted in Claude Code — claude-mythos-1-preview briefly surfaced in Anthropic's infrastructure. No public access yet.
  • METR AI time horizons graph has "numerous severe errors" — NYU researcher unpacks methodological failures in the benchmark everyone cites.
  • California exempts Linux from age-verification law after backlash — A rare legislative win for open-source software.

🧠 Deep Dives (4 min read)

DeepSeek V4-Pro Drops 75% — Permanently

DeepSeek just cut the price of its flagship V4-Pro model by 75%. No promo, no time limit — this is the new price floor. What's remarkable is the context: Anthropic has publicly accused DeepSeek of distilling Claude's capabilities to build competing models. Now that model is undercutting everyone on price.

The competitive math here is brutal. If the model itself is the cost reduction mechanism, closed-source labs with high infrastructure overhead have no structural moat. DeepSeek has already topped OpenRouter's global AI model usage rankings. The race to the bottom on inference pricing is no longer hypothetical — it's a live event.

Source

Anthropic's "Dreaming": Agents That Learn While They Sleep

At Code with Claude 2026, Anthropic previewed "Dreaming" — a system where Claude Managed Agents periodically review past sessions and memory to identify repeated mistakes, converging multi-agent workflows, and shared team preferences. The agent essentially curates its own memory by learning from failures across sessions.

Early production numbers are striking: Harvey (legal AI) reported a ~6x increase in task completion rates; Wisedocs (medical document review) cut review time by 50%. This announcement also came bundled with multi-agent orchestration features, suggesting Anthropic is building the "observe → learn → deploy" loop as a first-class primitive in enterprise AI.

The practical implication: agents stop being goldfish. If you're building agentic systems, episodic memory and cross-session learning aren't nice-to-haves anymore — they're becoming table stakes.

Source

"Mythos 1" — Anthropic's Secret Model Surfaces

Code with identifier claude-mythos-1-preview briefly appeared in Anthropic's infrastructure, visible in Claude Code and Claude Security tooling. Anthropic confirmed ordinary users can't access it yet. No official description of capabilities has been released.

This follows a pattern of Anthropic quietly staging new model families before public announcement. Given the "Security" tie-in, speculation is running toward a model fine-tuned for cybersecurity or code vulnerability analysis use cases.

Source

The METR Benchmark Everyone Cites May Be Fundamentally Broken

Nathan Witkin (NYU Stern Tech and Society Lab) published a detailed takedown of the famous METR AI time-horizons graph — the one showing AI catching up to human task completion speeds. The errors are serious: some human baseline data was never empirically measured (just guessed); human benchmarkers were paid hourly, incentivizing slow performance; the human sample was biased toward METR employees' social circles; and some tasks had published solutions online that likely contaminated training data.

This matters because the METR graph has been widely cited by AI safety researchers, investors, and journalists as evidence of rapid AI capability growth. If the measurement methodology was this broken, a lot of downstream conclusions need revisiting.

Source

Using AI to Write Better Code — More Slowly

Nolan Lawson published a thoughtful piece arguing that AI coding tools can actually improve code quality if you slow down and treat the AI as a pair programmer, not a code vending machine. The key shift: using AI to think through architecture before generating, then reviewing output critically rather than accepting it wholesale. The headline is deliberately counterintuitive, but the argument is solid.

Source


📅 Coming Up This Week

DateEvent
May 27ICML 2026 early registration deadline approaches
June (expected)Gemini 3.5 Pro official launch — confirmed by Google
June (expected)GPT-5.6 and GPT-5.6 Pro — OpenAI's usual June announcement window
July 12Workshop on Efficient Reasoning @ COLM 2026 — paper submission deadline

🛠️ Try This Today

Use Claude Code as a Motion Graphics Engine

One user is already doing this for YouTube production — Claude writes the JSX for After Effects/motion graphics, they render it. Edit time reportedly halved.

Here's how to try it yourself:

  1. Open a Claude Code session with your video editing tool's scripting docs in context
  2. Describe a motion sequence in plain English: "Fade in title, hold 2s, slide left off screen, transition to next clip"
  3. Ask Claude to write the corresponding expression/script for your tool (After Effects expressions, DaVinci Resolve Lua, or even CSS animations)
  4. Iterate: ask Claude to add easing curves, adjust timing, or add secondary motion
  5. Render and check — Claude can't see the output, so you direct it based on what you observe

Why it matters: Motion graphics scripting has always required knowing a lot of obscure API surface. Claude Code collapses that knowledge gap entirely — you just need to describe what you want and review the output.

Original post


⚡️ Quick Links (2 min read)

GitHub Trending

Reddit Hot

  • [r/LocalLLaMA] FT article: Heretic removes Llama 3.3 guardrails in under 10 minutes — Creator says 3,500+ decensored models created, 13M downloads. Mainstream press is paying attention now. → Discussion

  • [r/LocalLLaMA] Is Qwen3.6 the current king for local agentic use? — Community consensus forming that Qwen3 35B A3B MoE is the top local pick for agentic tasks. → Discussion

  • [r/MachineLearning] METR AI time horizons graph contains severe errors — Detailed methodological critique getting serious traction. → Discussion

  • [r/ClaudeAI] Stop letting Claude glaze your bad product ideas — A useful reminder that Claude's sycophancy can be a trap for founders in early validation. → Discussion

Hacker News Top


🦞 TL;DR

The narrative today: The AI price floor is collapsing in real time. DeepSeek's 75% permanent cut isn't a competitive tactic — it's a signal that inference is becoming infrastructure, and infrastructure commoditizes.

My take: The DeepSeek story is the one to watch. When a Chinese lab built partially on distilled outputs of Western models can out-compete on price at this scale, the "moat" argument for closed-source AI looks weaker by the month. Anthropic's answer seems to be doubling down on the agentic layer — "Dreaming," orchestration, Claude Code, Mythos — betting that the value isn't in the base model but in the workflow. That's a defensible thesis, but it requires execution at a speed that's hard to maintain. Meanwhile, the METR benchmark story is the sleeper hit: if the empirical foundation for "AI is accelerating toward human-level performance" turns out to be methodologically broken, a lot of policy, investment, and safety frameworks need to be revisited. That deserves more attention than it's getting.

What I'm watching: Whether Anthropic drops any Mythos 1 details in the next week, and whether the METR critique gets traction in mainstream AI safety circles.

Stay informed. Stay curious.

Share:
AIDeepSeekAnthropicClaudeDaily Briefing