AI Briefings·6 min read

AI Morning Briefing — October 5th, 2026

Lyubo
Lyubo·
AI Morning Briefing — October 5th, 2026

OpenAI sets April 2027 shutdown for gpt-5.1/5.3-codex/5.4-nano, DeepSeek Harness goes desktop, Strata runs a 125B MoE on one GPU, and RemoveMacAI strips Apple Intelligence.

AI Morning Briefing — October 5th, 2026

Your daily digest of what's happening in AI, straight from the trenches.


🚀 Headlines (30 sec read)

  • OpenAI sets an April 2027 shutdown for gpt-5.1, gpt-5.3-codex and gpt-5.4-nano — Replacements are gpt-6-sol and gpt-6-luna, and they are not all the same target
  • DeepSeek Harness ships desktop apps — The MIT-licensed agent runtime now installs on macOS and Windows, with scheduled tasks and background execution
  • Strata runs a 125B Qwen MoE on one consumer GPU — A Hacker News hit claiming ~94 tok/s on a 12GB card by keeping hot experts in VRAM
  • RemoveMacAI strips Apple Intelligence from macOS 27 — Disables the features and deletes the downloaded models (463 points on HN)

🧠 Deep Dives (4 min read)

OpenAI deprecates three API models, shutdown April 1, 2027

OpenAI's deprecations page has a new 2026-10-01 entry with six months' notice. gpt-5.3-codex and gpt-5.1 map to gpt-6-sol; gpt-5.4-nano maps to gpt-6-luna. The nano migration is the one to watch: if you bulk-replace every old ID with Sol, you change cost and latency for your cheap, high-volume calls. Using a new model in ChatGPT says nothing about your API code, so grep for hard-coded model IDs and diff outputs on real inputs before switching. → Source

DeepSeek Harness v0.2 gets a desktop app

DeepSeek's open-source agent runtime (MIT) now has official installers for Apple silicon Macs and Windows. Pick a local folder as a workspace, hand the agent a task, and it can run scheduled jobs that survive restarts and keep running after you close the window. Everything is a plugin, including the agent loop, and it is not locked to DeepSeek: it lists Anthropic, OpenAI, Moonshot and Z.ai providers plus custom endpoints. It is a developer preview with no security audit, and the terms warn that agents can execute code and commands with access to your files and credentials. With official DeepSeek models the session logs are collected; with custom models your inputs go straight to that provider. Run it in a scoped folder with minimal permissions. → Source

Strata: a 125B MoE on a gaming PC

Strata is an MIT-licensed inference engine built on llama.cpp/ggml for Qwen3.8-Flash-Next. It keeps the few thousand most-used experts in VRAM, holds all 24,576 experts in system RAM, and puts the rest on SSD. A small draft model proposes tokens and the big model verifies them in one pass. The README claims about 94 tok/s generation and 2,650 tok/s prompt processing on an RTX 5070 (12GB) at Q2_0. The HN title says RTX 4090 at 100 tok/s, so the numbers vary by card. Needs 12GB+ VRAM, 32GB RAM (64GB recommended) and ~80GB disk. These are the author's own benchmarks, and Q2-level quantization costs quality, so test on your own tasks. → Source

Turning Apple Intelligence off on macOS 27

RemoveMacAI is a script for Apple silicon Macs on macOS 27 that disables Siri, Writing Tools, Genmoji and Image Playground, deletes the downloaded on-device models, and blocks re-downloads. It has status and revert commands, and changes persist through macOS updates. It is installed with a curl | bash one-liner, so read install.sh first. → Source


📅 Coming Up This Week

DateEvent
Q4Gemini 4 is reported to be in post-training, expected before end of 2026; no date announced
Dec 11OpenAI API removal announced June 11 takes effect (check the deprecations page for which models)
Apr 1, 2027gpt-5.1, gpt-5.3-codex, gpt-5.4-nano removed from the API

🛠️ Try This Today

Audit your code for deprecated OpenAI model IDs

  1. Run grep -rnE "gpt-5\.(1|3-codex|4-nano)" . --exclude-dir=node_modules --exclude-dir=.git
  2. Move model IDs into one config value per use case
  3. Point a staging copy at gpt-6-sol (or gpt-6-luna for nano) and compare outputs on 20 real inputs

Why it matters: You have six months, but cost and behavior differences are easiest to find before a deadline.


⚡️ Quick Links (2 min read)

GitHub Trending

Reddit Hot

  • [r/LocalLLaMA] Micron CEO says memory supply will be much tighter in 2027 and 2028 than in 2026 — Bad news for anyone planning a local rig → Discussion
  • [r/LocalLLaMA] From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck — Home-cluster power limits → Discussion
  • [r/LocalLLaMA] GLM 5.3 flash got a 50%+ performance boost for dual DGX Spark users → Discussion

Hacker News Top


🦞 TL;DR

The narrative today: Models are getting retired on a schedule, agents are moving onto the desktop, and big models are being squeezed onto consumer GPUs.

My take: The OpenAI deprecation is the only item here with a deadline, and the nano-to-Luna split is the kind of detail that bites people who find-and-replace. DeepSeek Harness is interesting because it is model-agnostic and MIT, but a preview agent with file and credential access and no audit belongs in a sandbox. Strata's numbers are the author's own, so I'd wait for independent benchmarks.

What I'm watching: Whether Gemini 4 lands before year end, and whether Strata's numbers hold up outside the README.

Stay informed. Stay curious.

Share:
AIOpenAIDeepSeekLocal LLMDaily Briefing