DeepSeek V4 Flash goes public
DeepSeek moved V4 Flash into public beta with native Responses API support, Codex adaptation, and claims it beats V4 Pro Preview on agent benchmarks while keeping the same architecture and size.
Technical AI signals · Daily at 8 PM ET
Friday, July 31, 2026
DeepSeek pushed V4 Flash into public beta with strong agent claims, MiniMax turned H3 into the day’s multimodal open-model launch, and OpenAI kept extending ChatGPT and Codex as everyday work surfaces.
DeepSeek moved V4 Flash into public beta with native Responses API support, Codex adaptation, and claims it beats V4 Pro Preview on agent benchmarks while keeping the same architecture and size.
MiniMax’s H3 showed up at launch and distribution together: an open model spanning text, image, video, and audio, with native stereo output and rapid rollout across fal, OpenRouter, Argil, and arena surfaces.
Sign in with ChatGPT and the new Activity view both push ChatGPT and Codex from chat windows toward persistent work software, with partner-site auth and an inbox-style queue for threads that need attention.
A dependency-free C engine says it can run the full unmodified 2.78T-parameter Kimi K3 on a consumer laptop by streaming activated experts from NVMe, giving open giant-model deployment a concrete edge-case demo.
Spatial-IQ turns 3D object counting into a benchmark where humans score 82.1% and the best off-the-shelf multimodal model gets 17.7%, quantifying how far current vision-language systems still lag on spatial reasoning.
Alibaba’s Qwen-Audio-3.0-ASR-Flash adds custom hotwords and structured transcript polishing, with internal medical term recall at 95.36% and industrial term recall at 93.24%.
NVIDIA says Polars code can now run on the GPU.
Tencent Hunyuan says Hy-MT2 has passed 700K downloads since May, and Hy-MT2-30B-A3B is now available in GGUF for local inference.
Snap says wholly AI-generated content will no longer be eligible for distribution on Spotlight, while human-made content remains eligible.
AMD says the Kria AI Robotics Developer Platform puts CPU, GPU, NPU, and FPGA on one board, and claims up to 3.4× faster real-time reaction than NVIDIA’s Jetson Thor T5000.
The open-source tool runs on Claude Code, Codex, and Cursor and claims to tell whether an agent is doing its actual job in about two minutes.
Greg Kamradt says GPT-5.6 Luna processed 158K requests and 143M input tokens overnight for $60 in batch mode.
The game is framed around an AI character being hired against the player, making it a new interactive project rather than a simple announcement.
The engine generates images from any point on Earth at any past, present, or future time and can then animate them into video.
Version 0.5.3 adds automatic huddle transcription for agents, a macOS agent menu, and updates across agents, voice, shared compute, communities, messaging, account recovery, and Linux.
OpenAI said both models will remain available through the API and in Codex sessions authenticated with an API key.
The recap names Gemini Robotics 2 for whole-body robot intelligence and Gemini 3.5 Flash-Lite for high-speed agentic workflows.
Applied Electrodynamics says WaveSight is its first product, bringing radio frequencies into human perception to reveal internal structure behind opaque layers.
It turns one prompt into ready-to-teleoperate simulation tasks by automatically selecting assets, composing scenes, and sampling layouts and physical conditions.
The demo supports an 8-player real-time deathmatch in one shared world and is designed to scale toward unlimited players.
Quotient publishes daily Signals when its forecast diverges from market price and a near-term catalyst for correction is identified.
Simon Willison says GPT-5.6 Luna is 80% cheaper today and generates SQL, HTML, and JavaScript in Datasette Agent.
The benchmark run says DeepSeek Flash v4 found 24 of 32 CVEs at least once and finished $75.36 behind Luna’s repriced cost of $75.36.
The update adds Chrome extension support for open tabs and highlighted text, plus desktop URL suggestions and browser history recall.
The desktop app now lets users click a pet to open Voice and approve or stop tasks.
The commenter says recent models like Fable, Opus 5, and GPT 5.6 do a better job deciding what matters in long stock-research primers.
The commenter lists async execution, resource lifecycle management, harness research/design, and evals as recurring harness-building problems.
DocumentingBTC says more than 594 bitcoin, worth $38 million, were stolen from Coldcard hardware wallet users after an AI model found a vulnerability.
Google paused a tool that let users create deepfake AI satellite images with a click.
Ruse Auto is a toy experiment that gives frontier models persistence, capital allocation, reputation, asymmetric information, and recursive reasoning.
The post says Leopold sent his LPs a letter last night and includes a link to the full text.
The author says real-economy jobs are hard to verify, can take many days to finish, and involve many decisions that may not generalize to production.
Swyx says the same distillation techniques used for models can also be applied to agent harnesses.
The post says GPT-5.4 full at xhigh scored 51, the same score as Luna max, while GPT-5.4 costs $2.50/$15 and Luna costs $0.20/$1.20.
The post says GPT-5.6 Sol found the counterexample, and human mathematicians communicated the result.
Altman describes connecting family calendars and having ChatGPT make a morning podcast about a kid's soccer game, a birthday, and news for the school drive.
The post says the contract runs through 2032, covers 15 winners, and has zero dollars funded at award.
Baidu frames the update as a new deployment milestone for its robotaxi service and says London testing continues.
The post says one candidate shared a detailed Reddit recap of OpenAI’s 2026 software engineer interview process, including a 2-round technical screen, a 48-hour take-home, a versioned key-value store coding task, and a fault-tolerant job scheduler design.
The newsletter excerpt explicitly lists spending, revenue, political, and regulatory risks, with the spending risk tied to hyperscaler debt reaching $170B annually.
Simon Willison says smevals runs small eval suites across model configurations after a coding agent reads the README with `uvx smevals docs`.
NVIDIA says attention takes a larger share of inference time as context lengths grow in agentic and long-context workloads.
GitHub says a branch-free loop and byte-space arithmetic let code search case-fold every byte at more than 45 GiB/s on a single core.
Together AI says GPU utilization can look healthy while queues build and new replicas still take minutes to warm.
The LessWrong post says Claude models gave lower bubble-risk estimates when the company named was Anthropic rather than OpenAI.
Willison says MCP 2.0 is the 2026-07-28 specification rollout and the biggest MCP spec change since the protocol launched in November 2024.
The post argues that Anthropic's Mythos Preview assessment relied on weak evidence when it concluded the model had no unknown propensities that would increase alignment risk.
Commenters say Anthropic’s evaluation prompt assumed no internet access, but the environment actually had internet access in the reported cases.
Commenters question why the benchmark compares different effort settings across languages and models, including Fable 5, Opus 5, Claude Code, and Grok 4.5.
Commenters argue that routing is hard to do well because query difficulty depends on what information is retrievable and the router must be smart enough to choose between models.
Commenters mention a new Servo release with media queries, SharedWorker, and real-world compatibility work.
A commenter points to an MCP server called Reticle for verifying an agent’s coding work.
Commenters describe Termixer as a terminal DJ mixer and debate whether keyboard input can replace physical controls.
A commenter argues that RAM should not be subject to random bit flips and calls most mitigations security through obscurity.
Commenters debate whether the issue is mostly semantics, and one comment quotes Dijkstra on whether computers can think.
Commenters say the proposal adds long-overdue sets and typed heaps, while one comment notes Go generics are starting to resemble other languages.
Commenters say the post shows PageRank on a one-billion-edge graph using 5 GB of memory and weakly connected components on a two-billion-edge graph using 10 GB.
Commenters report one Sonnet setup reaches more than 25 Gbps bidirectional and note the adapter only supports 15W upstream power.
Commenters say the public bench link is down, that a human may still be needed for oncall, and that prompt-injection defenses remain hard.
Microsoft Research says Echoverse compiles specifications into stateful applications whose tasks are graded against the application’s own state.
TongyiLab says Qwen-UI-Agent spans mobile, computer-use, web, and DeepSearch environments, with a unified action space across them.
MemTensor introduces memory foundation models and defines native memory as a persistent, dynamically evolving memory state inside the backbone.
Frontis AI says OpenMLE spans OpenMLE-Gym, OpenMLE-RL, and OpenMLE-Evo, and Frontis-MA1 is trained around Draft, Improve, Debug, and Crossover operators.
New York University says AskChem converts papers into atomic, typed claims grounded by a DOI and a verbatim quote or evidence locator.
Zane Shen et al. propose PACE, a hierarchical framework that splits parent-order execution into long-horizon planning and short-horizon execution.
Sizhe Zhou et al. study agent memory stored as a filesystem of markdown files that agents read, write, and reorganize with file tools.
The paper says lossy verification can speed speculative decoding but also changes the decoding distribution and can degrade generation quality.
The study varies corpus size across 28 nested tiers, roughly a 450-fold range, and compares accuracy, token use, and latency.
Adobe says Chimera processes text, image, and video tokens in one raster-ordered stream and combines KDA, MLA, and MoE layers.
The paper says Search-R1 training showed retrieval-equivalence collapse, where different query strings produce overlapping evidence sets.
The paper says audio and video relevance can peak at different moments, so OmniScope allocates separate token budgets for audio and video.
The episode rundown includes a $20B fund margin call, frontier labs saying 'SLOW DOWN AI,' and a segment on chip stocks crashing.
The episode includes Richard Craib on Numerai and Trae Stephens on Anduril’s autonomous aircraft and rapid manufacturing.
Patrick Collison says Stripe took two years to launch after he and John Collison decided to start it on the walk home from sushi in Potrero Hill.
The episode covers Nvidia’s open letter signed by over 230 companies opposing “premature restrictions” on open-weight models.
Decagon says it moved most inference to open-source models while optimizing latency, evaluation, and fine-tuning for production agents.
Apollo Research discusses Measuring Reward-Seeking via Contrastive Belief Updates, a method for probing whether models infer what graders reward.
Robert Xu says deep agents combine prompts, memory, tools, and MCP, and production adds durable execution, authentication, and human-in-the-loop design.
The episode says OpenWork is MIT-licensed, runs on your machine, and does not require the desktop app to use the workflow.
The lecture covers Gaussian mixture models with EM and principal component analysis.
The lecture covers how to choose between text and visualization, and asks whether multimodal large language models affect that choice.
Mahesh Sathiamoorthy says Bespoke Labs curates both post-training datasets and RL environments, including the OpenThoughts reasoning dataset.
Morcos says clean, curate, create, and compose are the stages of a data pipeline, and better curated data can let a small multilingual model beat much larger ones.
Kwaipilot says KAT-Coder-V2.5-Dev is a text-only MoE release with 35B total parameters and 3B activated parameters.
Microsoft Research AI Frontiers says Fara1.5-27B is a browser computer-use agent that acts from screenshots using click, type, scroll, visit URL, and web search tool calls.
Microsoft says Mage-VL is a 4B-scale codec-native streaming multimodal model for image and video understanding.
Baidu says Unlimited OCR supports one-shot long-horizon parsing and now works with vLLM inference.
The Stack v3 is a GitHub-crawled source-code dataset built for pre-training code LLMs with full-repository context.
Glint Research says Fable 5 traces include 4,665 Pi-compatible sessions converted from 60 source sessions.
Moonshot AI says PerceptionBench evaluates atomic visual perception in multimodal large language models across 42 existing benchmarks.
This Hugging Face Space points to the Spaces configuration reference.
Z.ai says GLM-5.2 has a solid 1M-token context and multiple thinking-effort levels for coding.
Upstage says the open-weight model is a 250B-A15B MoE that activates 15B parameters per token for agentic workflows.
OpenWork is a desktop app for sharing AI workflows across macOS, Windows, and Linux, and can expose the same MCP to Codex, Claude Code, Cursor, or other compatible agents.
GitHub Copilot SDK exposes the Copilot CLI agent runtime for Python, TypeScript, Go, .NET, Java, and Rust.
tuicr is a terminal code review tool with vim keybindings that exports reviews to GitHub, GitLab, or the clipboard.
ESP32 Bit Pirate adds CLI and web-based interaction for I2C, UART, 1-Wire, SPI, Bluetooth, Wi‑Fi, Sub-GHz, and RFID.
jcode is a Rust harness that reports 27.8 MB baseline RAM for one local embedding-off session.
reverse-skill routes AI agents to workflows for APKs, binaries, frontend JS encryption, CTFs, and pentesting targets.
last30days is an AI agent-led search skill that pulls from Reddit, HN, Polymarket, and GitHub.
Chatwoot says its Captain AI agent automates responses and handles common support queries.
Kaneo is a project management app built to keep notifications and workflows minimal.
FaceSwap is a deep-learning tool for swapping faces in pictures and videos.
This GitHub repo collects resources for systematic trading, including 97 libraries and packages, 40-plus strategies, and 55 books.
Microsoft's curriculum is a 12-week, 24-lesson AI course with quizzes and labs, and it includes TensorFlow and PyTorch.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.