Top AI News Weekly

Week 41 · Sep 28 – Oct 4, 2026

Three Frontier Launches in Three Days, and a Week of Videos Written in Code. Anthropic released Claude Sonnet 5.5 on 28 September at Sonnet 5's price, and its own table puts it ahead of Opus 5.5 on Terminal-Bench 4.0. OpenAI followed at DevDay on 29 September with GPT-6.1 Sol, an Agents API and Dots, always-on ChatGPT agents with their own cloud computers; CopilotKit and an independent developer each published an open-source version of Dots within hours. Google answered on 30 September with Gemini 4 Argon, the first Gemini 4 model, released first to vetted cyber defenders. The second story came from Opus 5.5 itself: at least six repos this week turn it into a film studio, writing whole videos as p5.js, Canvas or Three.js code instead of calling a video model, from a 43-style film skill to a URL-to-launch-video service. Open weights had a strong week too, with IFM publishing K2-Horizon's 375B mixture-of-experts model along with its training data and checkpoints, and IQuest, BAAI and Ai2 releasing models for coding, self-improving agents and cited science reports. Speech moved on several fronts (Eleven v4, Microsoft's streaming transcription, and two compact on-device recognisers), and two papers used GPT and Claude models to help refute Grad's 59-year-old fusion-physics conjecture.

55 launches and research drops that matter for enterprise AI builders—curated, tagged, and ready for your next roadmap sync.

New drops

55

Unique sources

50

Key themes

Agents · Video · Language

language models

Language Models

LLMs, reasoning models, MoE architectures, and language understanding systems.

Frontier ModelGoogle DeepMind

Gemini 4 Argon

Google DeepMind's first Gemini 4 model, announced on 30 September by Koray Kavukcuoglu and aimed at long-horizon work in software engineering, legal and finance knowledge work, and cyber defence, including finding and patching vulnerabilities on its own. Google reports 77.9% on DeepSWE v1.1, 68% on CWE-bench v1 (tied first), 51.3% on AutomationBench and 91.7% on LVBench, all self-reported. Output limit rises to 1 million tokens from 64K. Introductory API pricing is $2 per million input and $10 per million output tokens, with cached input 95% off. Vetted cyber defenders get it first through the Fairwind Program; paid API customers and Google AI Ultra subscribers follow, with no general availability date yet. Follows the Gemini 3.8 models covered in weeks 39 and 40.

View release ↗
Claude Sonnet 5.5Anthropic

Claude Sonnet 5.5

Released 28 September, six days after Claude Opus 5.5 (covered here in week 40), at Sonnet 5's API price of $2 input and $10 output per million tokens, with $0.20 cache reads and $2.50 cache writes. Anthropic states over 30% faster output and up to 30% lower cost per task than Sonnet 5. Its own table shows 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5 at 66.4%, plus 55.5% on CursorBench 4.0 (Sonnet 5: 34.1%), 80.1% on OSWorld 2.1 (57.0%) and 1844 Elo on GDPval-AA v2.1 (1449), all vendor-reported. Ships with the same cybersecurity safeguards as Opus 5.5. Available on the Claude Platform, AWS, Google Cloud and Azure as claude-sonnet-5-5. Context window not stated on the page.

View release ↗
GPT-6.1 SolOpenAI

GPT-6.1 Sol

Released 29 September at DevDay, one week after GPT-6 Sol (covered here in week 40). OpenAI says it nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's token prices. API pricing stays at $2 input, $0.10 cached input and $10 output per million tokens, with a 1,050,000-token context and 128,000-token maximum output. Reported results: a tie with Astra on DeepSWE, a narrow lead over Claude Opus 5.5 on AutomationBench, and a factual error rate cut from 11.4% to 7.7% at low reasoning effort compared with GPT-6 Sol. Available in ChatGPT Work and Codex on paid plans and in the API as gpt-6.1-sol. openai.com blocked automated reading, so details come from TechCrunch, DataCamp and Latent Space coverage; benchmarks are vendor-reported.

View release ↗
Fully Open MoEInstitute of Foundation Models (MBZUAI)

K2-Horizon-375B-A23B

The flagship of IFM's K2 Horizon family: a sparse mixture-of-experts model with 375B total and 23B active parameters and a native 512K context. Trained on about 15T pretraining tokens, four context-extension stages, then reinforcement learning across five merged expert models. Weights, training data, code and intermediate checkpoints from every stage are public under Apache 2.0, and the model card's artifact table marked them all available as of 28 September after the family's 3 September launch with six sizes from 0.9B to 375B. IFM reports 70.2% on Terminal-Bench 2.1, 65.3% on Toolathlon Verified and 87.3% on GPQA Diamond, all vendor-reported against older comparison models. Serving recipes cover vLLM and SGLang on 8 GPUs; the technical report was still marked in progress.

View model ↗
Open Agentic Coding MoEIQuest

IQuest-Q1

An open-weights mixture-of-experts model from IQuest for agentic coding, reasoning and multi-step tool use, released 28 September: 320B total and about 15B active parameters, 88 layers, 256 experts with 8 active, mixed sliding-window and full attention, and a 524,288-token context. Text-only. Multi-token prediction runs recursively at inference for speculative decoding. Served through SGLang or a forked vLLM on 8-GPU nodes and evaluated with Claude Code and mini-SWE-agent on DeepSWE, Terminal-Bench 2.1, CyberGym and an in-house CLI benchmark, with scores only shown in a chart. The licence is a modified MIT that requires commercial products to display "IQuest-Q1". The team calls it early-stage. Follows IQuest Coder (week 1).

View repo ↗
Open Science Report ModelAllen Institute for AI (Ai2)

AstaBrief 8B

An open-weights 8B model from Ai2, announced 2 October, that turns a research question and retrieved paper excerpts into a cited scientific report. Fine-tuned on 47,000 real researcher queries, then preference-tuned with DPO on 6,000 pairs, with training data filtered for citation density. It now powers Fast mode in Ai2's Asta research assistant at about 51 seconds per report, against 179 seconds for the Claude-powered Thinking mode. Ai2 reports citation and answer precision competitive with Claude-based pipelines and DR Tulu, and says 23% of users who tried Fast mode kept using it; figures are Ai2's own. Weights on Hugging Face under Apache-2.0.

View release ↗
Offline Voice TranslatorGoogle Creative Lab

Gemma Translator

An offline voice translator from a small team at Google Creative Lab that runs Gemma 4 E2B on-device through LiteRT-LM on a Raspberry Pi 5 with 8GB of RAM. Two people face each other at a kiosk, each with their own lane and a rotating language picker; speech is transcribed with Moonshine, translated by Gemma and spoken back in the other person's language with moonshine-voice, with no internet needed after setup. It ships a React front end tuned for 480x320 displays, a Python API server, a one-command script that installs it as a Pi kiosk service, and STL files for a 3D-printed case. Built with Google Antigravity and marked as not an officially supported Google product. Apache-2.0, 1,586 stars, created 3 August.

View repo ↗
vision & image

Vision & Image

Image generation, editing, restoration, OCR, and visual understanding.

Layout-Controlled Image GenBlack Forest Labs

FLUX 3 Image

The image model in Black Forest Labs' FLUX 3 family, launched 1 October, following FLUX 3 (week 30) and FLUX 3 Action (week 40). It is built for precise control: elements are placed with JSON bounding boxes on a 0 to 1000 grid, up to ten reference images can be combined into one layout, and multi-step edits recolour, replace or move single elements while leaving other pixels unchanged. It renders natively at 2K and 4K, up to 5456x3072 (about 16.8 megapixels), with no separate upscaling step. Available through the BFL API and a web playground, with non-commercial and commercial weight licences listed for self-hosting; coverage disagrees on when public open weights arrive. No benchmark numbers on the page.

View release ↗
Multi-Turn Image EditingIdeogram

Ideogram 4.5

An image editing model released 30 September for multi-turn work. Ideogram says each edit changes only what is asked and copies untouched pixels straight from the source, which stops the artifacts, pixel shifts and colour drift that build up over repeated edits. Uses include colour and lighting changes, replacing or translating text, product photography, interiors, restoration, sketch-to-image, style reference, reframing and depth-to-image, with editing up to 4,016x6,024 pixels. Four quality modes at native 2K, reported at 0.8 to 22 cents per image, in Ideogram, the API and partners including Runway, Picsart, Pika and Leonardo. Open weights are announced as coming soon. Follows Ideogram 4.0 (week 23).

View release ↗
Encoder-Free Unified ModelNVIDIA and University of Waterloo

PixelUMM

A unified model that both understands and generates images and video from raw pixels with one decoder-only Transformer and no VAE or separate vision encoder. Images enter as 16x16 patches and video as four-frame tubelets, and text, clean pixels and noisy pixels share self-attention on a Qwen3-8B backbone with separate understanding and generation experts. It covers image and video understanding and generation, image-to-video and video editing, in 1.7B and 8B sizes. NVIDIA reports DocVQA 90.42, ChartQA 82.96, AI2D 80.12, MVBench 70.53 and Video-MME 57.33 for the 8B model, vendor-reported. Paper 29 September (arXiv 2609.38597), weights 1 October; the code says Apache-2.0 while the weights are tagged "other".

View project ↗
Sports Vision Demosjeremyipark

Vision Demos

Six practical computer-vision demos on sport and fitness footage: dance-sync scoring between dancers, chin-up counting with ascent and descent timing, bouldering hold detection and move sequencing, 3D climbing-route reconstruction from iPhone LiDAR depth, running cadence, foot-strike and knee analysis, and deadlift rep counting with back-form classification. They combine ViTPose-Plus-Large for pose, SAM 3.1 for segmentation and Gemma-4-26B-A4B-IT for classifying form, and each ships its own README and config template. No published accuracy figures. Apache-2.0, 585 stars, created 8 September and still pushed on 1 October.

View repo ↗
video & animation

Video & Animation

Video generation, editing, inpainting, and real-time rendering.

Code-Only Film Skilllemomo-ai

Lemo OPUSCAR

A Claude Code skill, released 26 September, that lets an agent direct, animate, score and mix short films entirely in code, with no video generation model or stock footage. It ships 43 film styles, each with a style prompt and a demo film made by Claude Opus 5.5, from ink wash, ukiyo-e and risograph to rubber-hose cartoon, 16-bit pixel RPG, silent film and low-poly 3D. The headline demo, OPUSCAR 98, runs 6 minutes 31 seconds and redraws all 98 Best Picture winners from 1927 to 2025, each in a fitting style. Built on Canvas, WebGL and Three.js with Node and FFmpeg, installed through the Claude plugin marketplace. MIT per the README, 900 stars.

View repo ↗
Claude Cartoon Starter KitJohnHeibel

Claude Animation Base

A starter kit for hand-painted cartoon animation with Claude, extracted from PDoomVideo, a 2 minute 36 second music video for "I'm Upping My P(doom)" that Claude Opus 5.5 wrote end to end in Claude Code using p5.js, p5.brush and parallel subagents (1,754 stars, no licence). You clone the kit, open it in Claude Code and ask for scenes shot by shot; the model writes p5.js and p5.brush code that renders headlessly to MP4 with Node and FFmpeg. It includes the Clawd character with views, eyes, mouths, hats, dances and 31 acted emotions, plus an animation guide that tells the model how to act and paint. MIT, 738 stars.

View repo ↗
URL-to-Launch-Videodiggerhq

LaunchVideo

Paste a URL or a prompt and get a 20 to 40 second launch video without any video generation model. Claude Opus 5.5 writes the film as one 1920x1080 HTML document, and a serverless agent on OpenComputer checks it for errors, renders every frame in headless Chromium, encodes with ffmpeg and stores the MP4. In URL mode the agent pulls the page's title, headings, main colours and Google Fonts. Ships with a Next.js front end, live at launchvideo.io, and renders at about real time. 266 stars, created 24 September; the repo has no licence file and needs an OpenComputer account.

View repo ↗
Music Video Generationmakevoid

Motion Graphics Music Video Skill

A Claude Code plugin for Opus 5.5 that makes a motion-graphics music video from a song and a creative prompt. The agent researches, writes a character and scene plan for approval, then generates in reviewed waves of sub-agents, calling image and animation models through fal.ai, with green-screen character animation, audio isolation for lip-sync and sound effects. Composition uses Node.js and p5.js, Python for analysis and mixing, and Swift Core Image for effects. The author estimates about $30 of fal credits and 3 million Opus tokens for a 3-minute video. Unlike the code-only skills this week, it drives external generation models. MIT, 123 stars, v0.2.2 on 28 September.

View repo ↗
Code-Rendered Video Skillstuzhechen2005

opus-video-skills

Claude Code skills that have Opus 5.5 make complete videos, one skill per visual style, with all imagery and music generated in code and no image, video or audio model. Frames render in headless Chrome and encode with ffmpeg. painted-animation draws watercolour cartoon shots with p5.js and p5.brush, cuts music videos to the measured beat and adds word-by-word karaoke subtitles; kinetic-reel combines kinetic typography with three.js WebGL layers and a score synthesised from the same timeline. Installs as plugins. MIT, 97 stars, created 26 September; template portions credit the Claude Animation Base author.

View repo ↗
Self-Reviewing Video EditorYeJe-cpu

SeeCut

An AI editor, created 25 September, that turns talking-head or AI-avatar footage into edited short-form videos with motion graphics and watches its own cut to improve it. A coding agent runs the pipeline while Gemini, via the Antigravity CLI, analyses the footage and reviews each render. It uses real webpage screenshots as on-screen evidence instead of invented B-roll, renders with HyperFrames and keeps a change only after comparing versions in pairs, for up to three passes. Optional steps add HeyGen avatars, ByteDance Doubao speech and export to editable JianYing projects. PolyForm Noncommercial 1.0.0 per the README, 210 stars, tested only on macOS.

View repo ↗
Open-Source Video EditorOpenCut-app

OpenCut

Since Week 21 2026: the open-source CapCut alternative replaced its codebase with a rewrite that targets web, desktop and mobile from one Rust core. The roadmap lists an editor API, a plugin-first architecture, an MCP server so AI agents can drive the editor, a headless mode for automation and batch rendering, and an in-editor scripting tab. The rewrite is still scaffolding, with a GPUI desktop shell and an FFmpeg build workflow added on 24 September, and none of those agent features have shipped yet. The working editor moved to opencut-app/opencut-classic and still runs opencut.app (last release v0.3.0, April). MIT, 91,684 stars.

View repo ↗
Video Prompting GuideGoogle Flow (Google Labs)

Creative Prompting with Gemini Omni in Flow

Google Flow's official prompting guide for Gemini Omni Flash, published 30 September as an X article with ten techniques and prompt templates: set high-level creative constraints instead of micro-managing; use start and end frames as anchors to reduce drift; tag image and video ingredients for characters, settings and style; adjust pacing after generation; transfer and mix styles; animate text; keep continuous scene rules; direct cuts and transitions; make narrow conversational edits that end with "keep everything else the same"; and direct second by second with timecodes such as [0-3s]. It also recommends negative keywords. No matching Google blog page exists, so the X article is the primary source.

View release ↗
4K Video RefinerNVIDIA Research

SoL-Refiner

From NVIDIA's SANA team (SANA-Video 2.0 was covered in week 30): a refiner that takes low-resolution output from any base video generator and lifts it to 4K in a single denoising step at the target resolution. It is trained with high-resolution continual training, frame-based RL post-training and one-step distribution-matching distillation, and supports bidirectional and streaming inference. NVIDIA reports cutting MiniMax H3 4K generation from 152.3 seconds to 5.64 seconds (27x), plus 1.60x on SANA-Video, 2.81x on Cosmos-Nano and 3.46x on Wan, on 24fps clips of 81 to 189 frames on H100 and GB200, all vendor-reported. Code in the NVlabs/Sana repo; paper arXiv 2609.37969, 29 September.

View project ↗
Few-Step DistillationByteDance and Texas A&M University

DMAD

Distribution Matching as Adversarial Distillation, from Texas A&M and ByteDance, posted 1 October (arXiv 2610.02188). It turns distribution matching into a classification problem: two discriminator heads on one shared backbone learn density ratios directly, so no extra auxiliary models are needed. The authors report FID 1.04 for one-step ImageNet-64 generation, 84.6% human preference over rCM on Wan2.1-T2V-14B and 79.1% preference over DMD2 on MiniMax-H3 (covered here in week 32), all self-reported. Code is on GitHub under Apache-2.0 and weights are on Hugging Face.

View project ↗
Video Diffusion DistillationByteDance and UC San Diego

PDMD

Projected Distribution Matching Distillation, from ByteDance and UC San Diego (arXiv 2609.35768, 28 September), makes Distribution Matching Distillation more stable for few-step video diffusion by removing the part of the critic's error that lines up with the gap between student and critic predictions. The authors say it is a one-line change to DMD with no extra loss, network, data or training stage. They report 83.73 at 4 function evaluations on Wan2.1-T2V-1.3B (DMD2: 83.44) and 83.17 on MiniMax-H3 at 4 evaluations (DMD: 82.76), and two-step generation also works; results are self-reported. Code on GitHub (ZeamoxWang/pdmd) and a 4-step checkpoint on Hugging Face; the repo has no licence file.

View project ↗
audio & speech

Audio & Speech

TTS, STT, voice cloning, music generation, and audio processing.

Expressive TTS ModelElevenLabs

Eleven v4 and v4 Turbo

ElevenLabs' new text-to-speech models, released 28 September: Eleven v4, its most expressive model, and Eleven v4 Turbo, a low-latency version for voice agents. Both use a new architecture that reads tone, pacing and emotion from text, and support inline audio tags such as [whispers], more than 90 languages (up from 70), voice cloning from 10 seconds of audio, professional voice clones and context stitching for long-form audio. ElevenLabs gives Turbo about 100 ms median inference latency and about 150 ms to first speech, and coverage reports v4 leading Artificial Analysis's Speech Arena at Elo 1319. Available through the API, SDKs and ElevenAgents, including the free tier. Follows ElevenLabs Music v2 (week 24).

View release ↗
MAI Model PlaygroundMicrosoft AI

MAI Playground

Since Week 24 2026: Microsoft's limited-preview playground now hosts MAI-Transcribe-2, MAI-Voice-2.1, MAI-Image-2.6 and the MAI-Thinking-1 reasoning model, plus a Chatter beta voice feature. On 1 October Microsoft added three speech models: MAI-Transcribe-2-Streaming, its first real-time transcription model, covering 60 languages with automatic detection at $0.54 per audio hour through year end; MAI-Voice-2.1 at $22 per million characters across 23 languages; and MAI-Voice-2.1-Flash at $15, with about 150 ms end-to-end latency. Microsoft says Transcribe-2-Streaming ranks first on Artificial Analysis, self-reported. Also on Foundry, OpenRouter and Vercel.

View release ↗
Tiny On-Device ASRCactus Compute

Whistle

Cactus's speech recognition model for small devices, released 2 October as the companion to Needle 3 (covered here in week 39). It is a 16.9MB model for 16kHz mono audio up to 30 seconds per pass, transcribing English, German, French, Spanish, Italian, Dutch and Polish, and it returns transcripts, word-level timestamps with probabilities and speech embeddings. Cactus reports first token in 11ms and 1,319 tokens a second on an Apple M4 Pro CPU, and says it beats Whisper Base and Moonshine Tiny v2 on word error rate across LibriSpeech, SPGISpeech and Earnings-22, without giving the figures. It shares Needle's C++ engine, so speech and language can ship in one binary. Apache-2.0 weights on Hugging Face.

View release ↗
Compressed English ASRFermion Research

Phonon-2

An open-weight English speech recognition model compressed from NVIDIA's Parakeet TDT 0.6B v3, with weights uploaded on 28 September. The encoder is ternary-quantised to about 2.1 bits per weight, shrinking the download from the teacher's 2.5GB to 164MB. Fermion reports a 5.21% average word error rate across seven Open ASR Leaderboard English test sets, down from Phonon-1's 6.56% at 415MB, and says it beats its own teacher on parliamentary and meeting speech. It transcribes an hour of audio in about 20 seconds on an M5 MacBook Air, all vendor-reported. CC-BY-4.0, with a PyPI CLI, Docker images and a macOS dictation app called Detta.

View release ↗
Bhojpuri Voice Cloninggentleman101

Bhojpuri F5-TTS

A rank-32 LoRA (10.1M of 347M parameters trainable) on AI4Bharat's IndicF5, trained on 9.4 hours of IISc's SYSPIN Bhojpuri corpus from two studio speakers, for zero-shot voice cloning from a 3 to 10 second reference clip. On a 32-sentence test set built around Bhojpuri-versus-Hindi pronunciation, pitch-contour correlation rose from 0.502 to 0.601 and pitch error fell from 3.76 to 3.21 semitones, improving all eight categories (self-reported, pitch proxies only, no native-speaker listening test). The author calls it a personal learning project, not production-ready, and asks users not to impersonate the two corpus speakers. Weights CC-BY-4.0, code MIT.

View model ↗
3d & spatial

3D & Spatial

3D generation, reconstruction, PBR materials, depth estimation, and spatial computing.

Explorable World ModelInSpatio

InSpatio-World 1.5

Follows InSpatio-World, covered here in week 31 as a 1.3B model that turns a reference video into an explorable 4D world. Version 1.5, with weights out on 24 September, accepts a single image, a set of four images, a panorama or a video, allows wider camera moves while staying consistent, and adds next-view prediction for dynamic video scenes. The checkpoint is still 1.3B parameters on the Wan2.1-T2V-1.3B architecture, with Depth Anything 3 for depth and camera estimation; you supply a target camera trajectory and video input is resized to 832x480. Code Apache-2.0, weights on Hugging Face and ModelScope, demo at world.inspatio.com. No speed or benchmark numbers for 1.5.

View project ↗
Agentic Real2SimXu, Bian, Wang, Sun and Gao

AHa-3D

A tool harness that has GPT-6 Astra (covered in week 36) turn an ordinary indoor video into an editable, interactive Blender scene in five stages: observe, build, verify, interact and deliver. Pi3X recovers cameras and metric geometry, SAM3 segments walls, floors and people, and TSDF fusion builds a reference mesh; a measurement tool sizes objects, RoomKit supplies 251 object and 252 material assets, X-ray overlays and Bullet physics checks catch misplaced objects, and GVHMR reconstructs human motion. The authors report better reconstructions than Astra alone at similar cost and time, shown through side-by-side examples rather than a numeric benchmark, and have only tested it with Astra. Apache-2.0 code, 249 stars, September 2026.

View project ↗
Dense 3D Point TrackingCarnegie Mellon University and Meta

TrackEverything

A 3D point tracker from CMU and Meta (arXiv, 24 September) that tracks every visible point in long videos, starting each track as soon as a point appears. It stores the video as persistent 3D scene tracks in world coordinates, so cost grows with scene geometry rather than video length; tracks are merged at window boundaries through voxels, a refiner separates static from dynamic points, and full trajectories are decoded only for dynamic ones. The authors say it is the first such tracker to handle videos over 1,000 frames within 40 GB of GPU memory, and report more than 20% better APD on TAPVid-3D than open-source dense 3D trackers, self-reported. A code repo exists while the page still says code is coming.

View project ↗
Human-Object InteractionUniversity of Tübingen and Max Planck Institute for Informatics

PAMI

Part Anchored Motion for Interaction generates full-body human-object interactions from text. It comes from Andreas Geiger's and Gerard Pons-Moll's groups at the Tübingen AI Center and the Max Planck Institute for Informatics. Body-part anchors vote for the object's motion, borrowing the idea from the Hough transform; PamiGen generates the interaction in a structured latent space and PamiRefiner keeps correcting contact geometry with long-range and short-range surface sensing. The authors report 14.5% higher contact recall than the previous best on the InterAct dataset, self-reported. Code on GitHub under MIT (repo created 29 September); weights not mentioned.

View project ↗
Prompted 3D Part SplittingCarnegie Mellon University

Point2Part

A CMU model that splits a whole 3D shape into parts in one pass, so parts never overlap or leave gaps, with users choosing parts through point prompts. One model handles single image to parts, mesh to parts and part segmentation. The authors report a Chamfer distance of 0.0534 for image to parts (OmniPart: 0.0643), 0.0273 Chamfer with F1@0.05 of 0.844 and fully watertight output for mesh to parts, and 69.80 segmentation mIoU, rising to 74.84 with four prompts, all self-reported. Code on GitHub (created 29 September) with no licence file; weights not mentioned.

View project ↗
Any-Skeleton MotionPrinceton, UC Berkeley, MIT and NTU

UniMate

Since Week 37 2026: preview checkpoints went up on Hugging Face on 27 September, after the training and inference code (6 September) and the UniML3D dataset and processing pipeline (30 August). UniMate generates skeletal motion from text for arbitrary skeleton types, from bipeds and quadrupeds to birds, sea creatures, insects, snakes and rigid objects, without retraining per skeleton. UniML3D holds 13,006 text-paired motion sequences. A preprocessing pipeline for new rigs is still on the to-do list. SIGGRAPH Asia 2026, MIT, 1,261 stars.

View repo ↗
agents & automation

Agents & Automation

Agent frameworks, orchestration, coding agents, desktop agents, and embodied AI.

DevDay RecapOpenAI

OpenAI DevDay 2026

OpenAI's 29 September developer day. Beyond GPT-6.1 Sol and the always-on Dots agents (both listed separately this week), it put an Agents API into public beta with hosted execution, memory, tools, multi-agent support and computer use. A Decisions API entered limited preview, using the Luna model for routing and classification. Codex gained cloud environments and Codex Security Cloud, and an Ultrafast tier that OpenAI says runs up to 8x faster in Codex and 6x in the API. openai.com blocked automated reading, so this recap rests on InfoQ and CNBC coverage; speed figures are OpenAI's own.

View release ↗
Always-On ChatGPT AgentsOpenAI

OpenAI Dots

Announced at DevDay on 29 September: always-on agents inside ChatGPT, each running on GPT-6 Astra (covered in week 36) with its own cloud computer and browser. A dot works toward goals you set while you are away, connects to more than 4,000 apps through plugins, can juggle several projects and can reach your laptop if permitted. You can message or call a dot in ChatGPT on desktop, web and mobile, or message it in Slack and Teams. Users set what a dot may do alone, what needs approval and what is off-limits. Included with Pro (outside the EEA, Switzerland and UK) and Business Premium, Enterprise in beta, and dot usage does not count toward plan limits at launch. Based on The Next Web, Latent Space and Drag coverage because openai.com blocked automated reading.

View release ↗
Open Agent Coworker TemplateCopilotKit

OpenDots

CopilotKit's open-source template for always-on AI coworkers, created on 29 September, the day OpenAI launched Dots, and following OpenMuse and OpenTag from the same team. Each Dot gets a name, role, instructions and permitted tools, plus an optional persistent computer with browser, files and terminal, per-Dot permissions and human takeover. You can text it, call it with realtime speech while another agent keeps working, or mention it in Slack. Approve-or-decline cards gate saves into Spaces, a nested document workspace. Built on AG-UI, the CopilotKit runtime and TanStack AI for web and mobile. Self-hosted only, alpha, MIT, 2,641 stars in five days.

View repo ↗
Self-Hosted Agent WorkspaceAnil-matcha

Open Dots

A second self-hosted, open-source take on OpenAI's Dots, independent of OpenAI and of CopilotKit's OpenDots. It offers assistant personas, streaming chat stored locally in SQLite with encrypted credentials, and a deny-by-default action gateway that pauses higher-risk actions for approval and logs audit events. Composio provides app connectors, /search adds web search, and an optional Docker and Playwright runtime gives it a computer. Bring your own inference endpoint. MIT, labelled an early prototype that is not production-ready. The repo's contents were replaced on 29 September and its star count belongs mostly to an older project that lived there, so it is not quoted.

View repo ↗
Personal Agent AppCopilotKit

OpenMuse

Moved from Week 40. A self-hostable personal agent app built on CopilotKit and AG-UI that gives the agent a persistent Chromium browser and an optional isolated Linux terminal for browsing, commands and files, with approvals and durable task plans, Gmail and Calendar over OAuth, PDF form filling and finance tracking. The front end is Expo and React Native for iOS, Android and web. Since last week it has added an E2B Desktop sandbox for the agent computer (2 October), optional Parallel web search, and answers to ambiguous requests as choices with sourced comparison cards. MIT, 3,861 stars (up from about 2,400), still alpha.

View repo ↗
Agent Hardware SDKMeta

Muse Gadgets

Open-source SDKs, released 2 October under Apache 2.0, for building physical devices that talk to Meta's Muse personal agent: an ESP32 firmware and device SDK reported to support 15 boards, from status lights to touch screens with push-to-talk, and a Linux SDK for Raspberry Pi and similar computers. Meta is also shipping Muse Home Link in October, a USB-C dongle that connects Muse to home Wi-Fi and smart devices such as TVs and speakers, free to US subscribers at one per account. The SDK repo is facebookincubator/muse-gadget-sdk. The board count comes from press coverage, not Meta's page.

View release ↗
ComfyUI Built-in AgentComfy Org

Comfy Agent

An agent built into ComfyUI, launched 1 October, that plans, builds and runs node workflows from plain-language requests and works on the canvas while you edit. It reads visual assets and existing workflows, fixes errors, compares techniques, runs up to five parallel chats with full history and supports public and private custom skills. In beta on Comfy Cloud with a free trial on existing Comfy Credits and token-based pricing; Comfy Desktop support is expected within weeks, local ComfyUI after that. The post names neither the underlying model nor per-token prices. Adds a native agent to the earlier Comfy MCP work (weeks 27 and 34), which let outside agents drive ComfyUI.

View release ↗
Self-Improving Agent ModelBeijing Academy of Artificial Intelligence (BAAI)

AREX-2

A 27B dense model from BAAI, uploaded 29 September and fine-tuned from Qwen3.8-27B, trained to improve a solution over several rounds by proposing, measuring, reflecting and revising. It learned on machine-learning and algorithmic coding tasks with verifiable feedback, and the authors say the skill carries over to deep research. BAAI reports 70.7 on Frontier-CS (Agent Track), 81.8 on MLE-Lite, 84.0 on BrowseComp, 92.2 on GAIA and 93.8 on DeepSearchQA, all self-reported in the model card. 262K-token context, Apache-2.0 weights and a companion paper (arXiv 2609.38288).

View model ↗
Agent Memory LayerAsymptote Labs

Beacon

An open-source memory and trace layer for AI coding agents, written in Go. It captures full session history across Claude Code, Cursor, Codex, OpenCode, Cline and more than 20 other harnesses, including prompts, tool calls, commands, edits, approvals, MCP activity and tokens, in one OpenTelemetry-based event model, then serves reviewed workflows, corrections and repo conventions back to future agents over MCP or Agent Skills. This week added beacon mcp connect to register its MCP server in harness configs, git commits linked to the agent sessions that wrote them, and threat rules v1.1 for spotting secret reads and indirect prompt injection. Local-first, though setup preselects the hosted option. MIT, 1,755 stars, v1.3.29 on 28 September.

View repo ↗
Spec-Driven Agent Codingopen-gsd

GSD Core

A meta-prompting and spec-driven development framework that runs AI coding agents in a five-step loop per phase: discuss, plan, execute, verify, ship. It fights context rot by pushing research, planning and execution into fresh-context subagents, each starting with a clean 200k-token window, and keeps state across sessions in files such as STATE.md and CONTEXT.md. Installs with npx @opengsd/gsd-core into Claude Code, OpenCode, Codex, Copilot, Cursor, Windsurf, Kimi CLI and others. MIT, 10,139 stars, v1.15.0 on 26 September after weekly releases since August.

View repo ↗
Spec-Driven Coding Agentopen-gsd

GSD Pi

A local-first terminal coding agent from the GSD Core team for long, autonomous, spec-driven work. Work is broken into milestones, slices and tasks that auto mode plans, implements, verifies and advances, with changes isolated in Git worktrees and requirements, decisions, plans and validation evidence kept in a .gsd/ folder. Models route through API keys, OAuth, or external CLIs such as Claude Code and Cursor Agent. Terminal UI by default, a web control plane via gsd --web, install with npx @opengsd/gsd-pi. MIT, 1,287 stars; latest tagged release v1.20.1 on 18 September, with commits still landing.

View repo ↗
Request-to-PR AgentThe01Geek

PRFlow

A Claude Code plugin that takes one feature request to a review-ready pull request in four phases: a spec written into a GitHub issue, implementation with tests, a review-and-fix loop that keeps fixing and re-reviewing until it approves, and docs. A separate shadow pass then re-checks the approval, and a weekly retrospective reads merged PRs and proposes improvements to the project's skill extensions. Runs locally with zero config or on GitHub Actions. v2.56.0 on 1 October makes reviews re-verify earlier defects before rejecting and records disputed criteria in the PR description. MIT, 99 stars; effectiveness claims are the author's own.

View repo ↗
Humanoid Tactile LocomotionTsinghua University

TactileStep

A Tsinghua University framework and CoRL 2026 spotlight that adds sole pressure sensing to humanoid walking control. A lightweight tactile simulator turns rigid foot contacts into a pressure array, the policy learns from normal force, contact area and centre of pressure, and rewards that depend on gait phase encourage softer landings. On a real Unitree G1 the authors report 48.8% lower impact force on platform ascent, 30.1 dB lower peak noise on stair descent and 23.8% more contact area, with success rates held or improved, all self-reported. Code is listed as coming soon (arXiv 2609.28959).

View project ↗
Humanoid BadmintonTsinghua University, CUHK, Zhejiang University, BUCEA and DeepCybo

Humanoid Badminton

A CoRL 2026 paper (arXiv 2609.31840, 25 September) that teaches a humanoid robot badminton from limited human motion data with three-stage hierarchical reinforcement learning. Motion augmentation turns a few hitting clips into many stroke variations, a high-level planner picks skills from the shuttle's position, and an adversarial regularizer keeps movement natural. The authors say it is the first real-world humanoid racket-sport system to hold multi-skill rallies with human players, including forehand, backhand and jump returns. The page names no robot, gives no numbers and mentions no code release.

View project ↗
developer tools

Developer Tools

SDKs, inference engines, data pipelines, monitoring, and AI infrastructure.

Opus 5.5 Video Promptsyihui-dev

Awesome Opus 5.5 Videos

A collection of 475 prompts behind viral videos people made with Claude Opus 5.5 (covered in week 40), mostly Canvas, SVG, Three.js and GLSL animations where the model writes the code instead of using a video model. Each entry links the creator's original post, a preview image and a prompt file meant to be pasted straight into Opus 5.5, and the set is also published as a JSON dataset. Featured categories include motion graphics (58), explainers (16), 3D scenes (14) and games (12). Remakes play side by side on Skillry, the agent-skills library that runs the repo. MIT, 1,627 stars in about a week.

View repo ↗
Open MoE Training StackAllen Institute for AI (Ai2)

Olmo-core 3

Version 3 of Ai2's open training library, released 1 October with a rebuilt mixture-of-experts system that scales to trillion-parameter models. Ai2 reports going from 8 to 128 experts at about 3.2B active parameters (total capacity from 4.6B to 47B) with under 5% throughput loss, and 52,000 tokens per second per GPU on eight NVIDIA B300s against 19,400 before, about 2.7x faster. MXFP8 adds about 21% over BF16, and it reaches 858 TFLOP/s per GPU at trillion-parameter scale on 512 GPUs; throughput figures are Ai2's own. Apache-2.0, 1,664 stars. Follows the OLMo-core repo covered in week 14.

View release ↗
Codebase Diagram Generatortt-a1i

Archify 3.0

Since Week 36 2026: Archify, an agent skill for Cursor, Claude Code, Codex CLI and OpenCode that turns code or a system description into architecture, workflow, sequence, data-flow and lifecycle diagrams as self-contained HTML, reached 3.0.0 on 28 September. The release brings a redesigned reader that fits the diagram to the first screen, better routing for fan-out and return paths, a Lifecycle v2 lane layout, sequence legends that no longer cross lifelines and verified delivery across all five diagram types, while removing guided story views and share-card exports. 3.0.1 the same day adds update reminders. MIT, 76.8k stars.

View repo ↗
Agent Design Skillkaankiziltug

Logo Design Skill

An Agent Skills package, launched 26 September, that has Claude, Gemini CLI, Codex, Cursor or Copilot work through logo design from brief to production SVG. The agent writes 8 to 12 concept one-liners, builds three in SVG, tests them, and stops at a checkpoint before producing colour, lockups, a presentation board, an icon set and guidelines. It includes a library of more than 1,400 real-world SVG logos classified by mark type, geometry, typography and industry, plus dependency-free Python tools that audit SVGs and run 16px, one-colour, reversed and competitor-shelf tests. Installs as a Claude Code plugin. MIT, 1,718 stars, v1.4.3 on 27 September.

View repo ↗
YouTube Channel SkillsJakeschincariol

YouTube Agent Skill

Eleven Claude skills for running a YouTube channel: scripts built from 21 hook formulas with the hook scored first, title and thumbnail checked together as one pairing, timecoded edit decision lists from a transcript, retention-export analysis that finds where viewers leave, Shorts pulled from long videos, niche trend ranking by each video's multiple over its channel median, comment triage, SEO, chapters checked against YouTube's rules, weekly planning and a channel audit. Six call dependency-free Python tools. It writes but never publishes. Installs as a Claude Code plugin. MIT, 436 stars, single commit dated 16 September.

View repo ↗
research & safety

Research & Safety

Research papers, alignment, red-teaming, interpretability, and benchmarks.

LLM Recommendation RankerNetflix

GenRec

Netflix's write-up of GenRec, a recommendation ranker built on an in-house foundation LLM, with the paper on arXiv as 2608.10257 (August). Member history and context are written out as natural-language prompts instead of thousands of engineered features. Phase one adapts an open-source LLM to Netflix catalogue and behaviour data; phase two post-trains it with ranking labels and reward signals. It is served on vLLM in prefill-only mode, scoring the whole candidate set in one forward pass with no decoding. Write-ups of the post report results comparable to or better than the production ranker with about 40x fewer labelled examples, context cut from about 5,000 to 1,700 tokens, and significant gains in an A/B test. The base model is not disclosed, and the Netflix blog could not be read directly.

View release ↗
AI-Discovered PhysicsMatt Landreman; Gómez-Serrano, Liehr and Taylor

Grad's Conjecture Disproved with AI Help

Two papers posted 21 and 22 September refute Grad's 1967 conjecture that magnetic plasma equilibria in a torus must have reflection, axial or helical symmetry, a question central to fusion reactor design. Matt Landreman's paper (arXiv 2609.26742) gives exact analytic 3D solutions in elementary functions, one family with uniform rotational transform and one with sheared, and his repo says GPT-6 Astra Pro (covered in week 36) found them, with the original prompt and chat included alongside DESC checks and verification notebooks (MIT). Javier Gómez-Serrano, Lukas Liehr and Mitchell Taylor (arXiv 2609.24739) build families with only N-fold rotational symmetry and verify the proof in Lean 4; coverage says GPT and Claude models helped with calculations and Lean checking while the authors designed and checked the proof.

View repo ↗
AI Cipher DecryptionCarter Church

The Letter to Marmont

A write-up, published 18 September, of using GPT-6 Astra to decipher a 1,300-symbol homophonic cipher letter to General Marmont that had been listed as an unsolved historical cipher. Working from one low-resolution plate in a 1969 French army journal, the model spent about six hours transcribing the hand-drawn signs, telling pen variants apart and recovering the key by simulated annealing against French letter statistics. The post rebuilds the full 155-sign cipher table, gives French and English translations, redates the letter from 1807 to March 1809 and publishes verification scripts. According to the post, Satoshi Tomokiyo's Cryptiana catalogue now lists the cipher as solved. A personal blog rather than an institutional publication.

View release ↗