AI Morning Briefing — September 15th, 2026

Claude Fable 5.1 cracks a 370-year-old cipher, Anthropic launches Claude for Financial Advisors, and a mathematician proposes rebuilding math PhDs for the AI era.
AI Morning Briefing — September 15th, 2026
Your daily digest of what's happening in AI, straight from the trenches.
🚀 Headlines (30 sec read)
- Claude Fable 5.1 cracked a 370-year-old Royalist cipher that stumped cryptographers for over a century — 44 minutes, 176k tokens, zero human hints, and the "key" turned out to be the book itself.
- Anthropic launches Claude for Financial Advisors, wiring Claude straight into Schwab, BlackRock, and Addepar — client positions, cost basis, and CRM notes, all inside one chat.
- A mathematician says AI can already produce "a PhD thesis one hasn't even read" — and wants to blow up how math PhDs get evaluated — oral defenses over papers, before the credential stops meaning anything.
New from Cole Medin: "My NEW FAVORITE Skill — Claude Code Drives My Whole Computer" — a 400-line skill that hands any coding agent full screen control with zero extra tooling.
🧠 Deep Dives (4 min read)
Claude Fable 5.1 Solves a 370-Year-Old Royalist Cipher
The Cyphral Distich — two lines of 32 numbers apiece, appended to Sir Thomas Urquhart's 1653 book Logopandecteision — had sat unsolved since it was first posed as an open problem in Notes and Queries in 1899, later landing on cryptography historian Klaus Schmeh's list of history's top 50 unsolved encrypted messages. Claude Fable 5.1 solved it in 44 minutes and 176k tokens, with zero human interjections, according to vals.ai. The breakthrough insight: Urquhart's own text kept emphasizing the number 32, matching the cipher's structure, and the accompanying poem referenced "his own heart's wishes, and the Author's minde" — a hint that the key wasn't some external cipher alphabet but the book itself. For each of the 64 numbers, the model treated it as a word index into the corresponding one of Urquhart's 32 "Proquiritations" (verses), then took the first letter of that word. The plaintext: "O GOD UPHOLD KING CHARLS THE SECOND AND MAKE HIM THE SUPREME RULER OF THIS LAND" — fitting for a committed Royalist writing in the depths of the English Civil War. It's a small, delightful result, but a real one: sustained, creative, multi-hour reasoning over an archival puzzle with no training signal to lean on, not a benchmark trained for the occasion. → Source
Anthropic Launches Claude for Financial Advisors
Anthropic shipped a plugin connecting Claude to the software wealth advisors already use: Charles Schwab Advisor Services (balances, positions, transactions, cost basis, money-movement alerts), BlackRock's Advisor Center (model portfolios and institutional analytics), Addepar (governed portfolio data across public and private markets), plus Orion, Envestnet, iCapital, Wealthbox, Wealth.com, and Zocks — advisors pick which to connect during setup. Bundled "skills" cover the actual workflow of the job: pre-meeting prep, post-meeting notes and follow-up, portfolio rebalance review, compliance and AI-policy review, estate and tax briefings, and prospect intake. This is a different bet than most AI-in-finance pitches: instead of a chatbot bolted onto a dashboard, it's Claude sitting inside the systems of record — reading actual client holdings and CRM notes rather than summaries a human typed up first. That's also exactly where the stakes are highest; "money-movement status" and cost-basis data are not the place to discover a hallucination the hard way, and Anthropic's own pitch leans entirely on the custodial integrations doing the accuracy work, not the model. → Source
A Mathematician's Blueprint for Surviving the AI Math Boom
Mathematician Daniel Litt (University of Toronto) published an essay arguing academic math needs to restructure now, not after the disruption finishes landing. His starting claim: a year ago, internal models at OpenAI and DeepMind already scored the equivalent of an IMO gold medal, and today's systems can autonomously resolve major open questions — meaning, bluntly, "it's now possible to produce a PhD thesis one hasn't even read." Where AI still lags, per Litt, is theory-building, asking good questions, and exposition — the connective tissue of a field, not the individual proofs. His proposed fixes all point the same direction: stop grading the artifact, start grading the understanding. Replace thesis submission with a rigorous oral defense where the student explains the work until examiners are satisfied, regardless of who or what helped produce it. Weight hiring and admissions toward talks and sustained discussion over papers. Build seminar cultures where speakers must actually satisfy a room, not just circulate a PDF. It's a rare example of an academic proposing structural change instead of just flagging the problem — worth watching whether any math department actually adopts a version of it. → Source
New from YouTube (2 min read)
Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights — IndyDevDan
Covers: Why a single leaderboard like the Artificial Analysis Index hides more than it shows, and the five benchmarks IndyDevDan actually checks to pick models on performance, cost, and speed together instead of one blended score.
Example: On Terminal Bench, he shows GPT-6 Astra topping Claude Fable 5.1 on raw score while running roughly 4x cheaper per task on combined input/output tokens — a gap no single index number would ever surface.
→ Watch
My NEW FAVORITE Skill — Claude Code Drives My Whole Computer — Cole Medin
Covers: A lightweight "drivescreen" skill — under 400 lines, no extra tooling installed — that lets Claude Code or Codex control your screen through native OS commands (PowerShell, AppleScript) instead of a bulky computer-use harness.
Example: Medin uses it to run his entire morning setup (task manager, Obsidian, Docker containers) and once had it research a GitHub repo, install the app, and screen-test its features end to end while he worked on a different machine.
→ Watch
GPT-6 Built a City Out of Text — Matthew Berman
Covers: A quick demo of GPT-6 Astra generating a fully walkable 3D city rendered entirely in ASCII characters.
Example: Two prompts produced "After Hours" — a rain-soaked, infinitely-generating city full of pedestrians that Berman walks through live on screen.
→ Watch
GPT-6 Made a Fall Guys Game — Matthew Berman
Covers: GPT-6 Astra building a playable Fall Guys-style obstacle course game, complete with sound effects and physics.
Example: One prompt produced a working game; a single follow-up feedback prompt refined it into the version Berman dares viewers to beat — his best time is 31 seconds.
→ Watch
📅 Coming Up This Week
| Date | Event |
|---|---|
| Sept 29 – Oct 1 | The AI Conference 2026 lands in San Francisco, with dedicated tracks for AGI, LLMs, and agentic AI |
| Ongoing | Clay Institute mathematicians and independent researchers continue picking apart OpenAI's contested Navier-Stokes proof |
| This week | Cryptography historians start trying to independently verify Claude Fable 5.1's Cyphral Distich solution |
🛠️ Try This Today
Build Your Own 5-Benchmark Scorecard
Inspired by today's IndyDevDan Deep Dive — stop trusting one blended leaderboard number for model choice:
- List the specific tasks you actually lean on AI agents for this week — coding, research, ops, whatever.
- For each, pick one benchmark that measures it directly instead of a catch-all index: Terminal Bench for agentic coding, Apex Agents for knowledge work, Automation Bench for tool-using workflows.
- For your top two model candidates, compare all three axes together — score, cost per task, and speed — never score alone.
- Re-run the comparison monthly. Rankings move fast enough that a two-month-old scorecard can point you at the wrong model.
Why it matters: the model that tops a general index can lose badly on cost or speed for your specific workload — the only way to know is to check the benchmark that actually matches what you're building.
⚡️ Quick Links (2 min read)
GitHub Trending
- JustVugg/colibri — run frontier MoE models on hardware you already own, pure C, zero dependencies
- alibaba/open-code-review — a hybrid code review tool combining deterministic pipelines with LLM agents across languages
- debpalash/VoiceStudio — open-source voice cloning and audio creation supporting 646 languages
Reddit Hot
- [r/LocalLLaMA] "DeepSeek engineer reflections on RSI — burying my talent to yesterday" — a DeepSeek attention-kernel engineer's raw essay on watching AI close in on the work he loves, and choosing to stay at the company racing to open-source it anyway → Discussion
- [r/ClaudeAI] "Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows" — 571 upvotes on a MacRumors report that Siri's backend was built to be provider-agnostic from the start → Discussion
Hacker News Top
- Pion, an agent designed to run any company autonomously (363⬆️) — Andon Labs hands AI agents email, phone, banking, and a browser to actually operate a retail store and café, not simulate one
- "Dario, Please" (450⬆️) — the sharpest pushback yet on Monday's Amodei pacing proposal, calling it regulatory capture dressed up as safety
- GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review? (141⬆️) — a head-to-head on whether the cheaper model holds up on real review tasks
🦞 TL;DR
The narrative today: a 370-year-old cipher fell to a chatbot, a mathematician published an actual plan for adapting his field instead of just worrying about it out loud, and Wall Street's advisors got their own Claude integration — three very different speeds of institution meeting the same wave.
My take: the cipher is the most fun and least consequential story here — a genuinely cool party trick that doesn't change anyone's job tomorrow. Litt's essay is the one I'd bet on mattering most in five years, precisely because it's not a reaction, it's a redesign: grade the understanding, not the artifact, before the artifact stops meaning anything. And Claude for Financial Advisors is the one with the nearest-term teeth — it's not a demo, it's Claude reading real cost-basis data and money-movement status inside systems people's retirement accounts depend on, which is exactly where "the model was 95% right" isn't good enough.
What I'm watching: whether cryptography historians independently confirm the Cyphral Distich solve the way the Clay Institute is being asked to for Navier-Stokes, and whether Anthropic responds on the record to "Dario, Please" or lets it ride out the news cycle like OpenAI did with the RubyGems disclosure gap.
Stay informed. Stay curious.
Related Posts
AI Morning Briefing — August 26th, 2026
Anthropic overtakes OpenAI in quarterly revenue, OpenAI's Jalapeño chip beats Nvidia Blackwell, Claude's memory unifies across Chat and Cowork, plus DevDay Astra speculation.
AI Morning Briefing — June 20th, 2026
Nobel winner John Jumper joins Anthropic, Fable 5 stays #1 despite US ban, and Chinese AI seizes 60% of open-source API market
AI Morning Briefing — June 19th, 2026
US blocks Claude Fable 5 globally; SpaceX acquires Cursor for $60B; GLM-5.2 beats GPT-5.5 in agentic evals; ChatGPT drops below 50% market share