OpenAI pushes math claims into artifacts
OpenAI says Astra produced ten advances in math and theoretical CS, and the notable change is not just the claim but the release of Lean certificates and walkthroughs that let others inspect the work.
Technical AI signals · Daily at 8 PM ET
Saturday, August 1, 2026
Today split between OpenAI’s Lean-backed open-problem proof claim, DeepSeek turning long-context caching into a real latency win, and agent infrastructure spreading across memory, payments, and open runtimes.
OpenAI says Astra produced ten advances in math and theoretical CS, and the notable change is not just the claim but the release of Lean certificates and walkthroughs that let others inspect the work.
DeepSeek says disk-backed context caching cuts first-token latency from 13 seconds to 500 ms on high-reference 128K prompts, which is a concrete serving improvement for repeated long-context workloads.
The stack is getting more operational at once: GitHub opened Copilot’s agent runtime as an SDK, ByteDance rewrote DeerFlow around sub-agents and sandboxes, and DeepSeek is actively recruiting harness builders for internal testing.
Fresh repos and papers all point the same way: memory is moving from chat history hacks to explicit services and model components for long-term recall, reconstruction, and multi-agent reliability.
MoonPay’s Paybox turns bookings, shopping, token swaps, and prediction-market trades into callable agent actions, pushing agent UX closer to real transaction execution instead of read-only assistance.
Wonder combines camera-control cues, sparse retrieval over full-fidelity history, and a distilled student model to reach minute-scale interactive world exploration at 16 FPS.
Tau Robotics says human operators remotely control the robots while AI handles navigation, and the service is priced at $30 an hour.
One list of claimed Codex abilities includes speeding up footage, trimming clips, converting MOV to MP4, extracting audio, and making vertical social-media versions.
The post says Google has started pre-training a new foundation model, while Gemini 3.5 Pro stays in testing with trusted partners.
The WSJ says OpenAI fell behind Anthropic and is now trying to recover.
MiniMax H3 is now available through Vercel's AI Gateway.
The poster says Claude Fable 5 used 1.4 million tokens over 2 hours to generate a humanoid robot design, and shared the repo.
Gergely Orosz argues stacked diffs have less relevance because AI models now write more code, making GitHub look years behind.
Levels says Photo AI can generate videos with you or trained models directly inside the editor.
The post says Transformers use quadratic $O(L^2)$ memory and that Google’s method could avoid that scaling as context windows reach millions of tokens.
Rocket Lab says NITE-STAR is a $981M IDIQ program for space test and training infrastructure.
Emollick says the P&G study found AI blurred the lines between jobs, and OpenAI reported a similar finding.
The candidate contains only a bare URL with no extractable technical details, so it cannot support a substantive headline.
The post contrasts 2025 bullish AI-winner framing for $PLTR with a 2026 view that its fundamentals are catching up to valuation.
The commenter says many OpenAI employees hook ChatGPT to Slack, and coworkers dislike when a coworker's ChatGPT asks them for help.
The report says Twitch and Amazon may announce stream training for generative AI in coming months, with creators able to opt out of some data.
Action Model says selected ambassadors get early access to new products and campaigns plus exclusive community roles.
The video is about a TikTok challenge that the uploader says is destroying AI flock cameras.
Gemini says ChatGPT did not solve the 10 open problems and that doing so would signal superhuman research capabilities in pure mathematics.
The post says the team assessed Inkling and argues access should widen in stages.
The poster says they will join NYU as an Assistant Professor next fall and spend a year at @physical_int first.
The post offers only the line "what a privilege it is to be tested."
Chamath says he acquired almost 6GW of grid power and behind-the-meter capacity through 2029.
The Science paper describes a general-purpose biomedical AI agent designed to automate biomedical research workflows, including basic research and translation.
The post only links to a t.co URL and provides no additional details.
Simon Willison says the 304B-parameter DeepSeek-V4-Flash-0731 costs $0.14 per million input tokens and $0.27 per million output tokens.
datasette-apps 0.2a0 adds app_debug(), which opens an app invisibly in a 0-opacity iframe and runs JavaScript to test it.
The post summarizes a paper on the complexity of infinite-width networks and when functions can be learned with polynomial sample complexity.
Epoch AI says the brief covers FrontierMath, parallelizability and technological singularity, AI energy use, and signs of AI uplift.
The post argues technical AI safety researchers should pay more attention to RLVR trends, citing GRPO from the R1 paper 1.5 years ago.
The author describes experiments with a scientific agent called 5.6-sol and says the work was part of a Templeton-funded Proofs & Reasons project.
The post describes using AI to transcribe handwritten worksheets and diagrams from notes and reviews to extract patterns and forgotten insights.
The post proposes an RLVR approach that rewards models for red-teaming the training environment to reduce reward hacking.
Commenters say Seedance 2.5 looks unusually good for video generation, with one top comment calling the washing-machine ad example as strong as social media content.
Commenters say the practical consequence is that independent kernel checking still works, but users need current versions of both implementations.
Commenters discuss a nearly 800-page book on 64-bit assembly and programming.
Commenters point to NetBSD 11.0 release details, including npf(7) layer 2 and user/group filtering and a new x86 MICROVM kernel that can boot in about 10 ms.
Commenters refer to a kernel patch that mentions a bad AI-generated analysis in a ripgrep bug report.
Commenters say cloning a template database saves a lot of time in Postgres tests and is simple to implement.
Commenters compare Flint with Vega-Lite and say Flint is better for predetermined chat types but less flexible for custom work.
Commenters say the writeup uses Claude-style phrasing and question whether crawling Hugging Face was easier than crawling the broader web.
Commenters say the paper combines older winner-take-all ideas for K-modal generative models with modern diffusion and flow pipelines, and one commenter notes K-1 extra forward passes.
Commenters report that Solid Queue 1.6.0 adds fiber workers for Ruby and Rails job execution.
Commenters say the guide covers attention decode on AMD MI450 GPUs and relies on HIP and ROCm terminology.
Commenters say the repo claims a Linux-on-ESP32-31 setup that appears to boot enough to log in and run commands.
Commenters cite Censys ARC data showing 4,148 internet-exposed EtherNet/IP hosts that self-identify as Rockwell Automation or Allen-Bradley, with 2,945 in the United States.
The paper introduces Explorative Modeling, which explores K candidate matches between model generations and data.
SpatialCLI teaches VLMs to use specialist spatial tools and then internalize those capabilities.
PhiZero learns physical language from videos and uses a reason-then-render setup to predict future world evolution.
INTACT turns action-labeled, reward-free trajectories into an intent-to-action interface without test-time search.
The paper studies whether misleading knowledge in open information environments can push Deep Research agents to false conclusions.
Beacon frames tool use around mode adaptiveness and tool effect for multimodal reasoning tasks.
ShadowDancer learns unified dynamics representations from a video and its shadow to support frame-level control across actions.
Simile says it has raised $300 million total, including a $200 million Series B at a $2 billion valuation.
David Brumley describes training models through a ladder from crashing targets to reading and writing arbitrary memory and full exploits.
The episode says Theta Software builds environments and verifiers to measure long-horizon tasks more honestly.
The tutorial shows how to connect SuperCode CLI to Unreal Engine 5.8 using Model Context Protocol commands.
The episode says Poolside trained a model from scratch in eight weeks that beats a model eight times its size on a coding benchmark.
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash and uses speculative decoding with the same model structure as DeepSeek-V4-Flash-DSpark.
Kimi K3 is a 2.8T-parameter open-weight multimodal model with a 1-million-token context window and native vision.
Unlimited OCR Works is Baidu's long-horizon parsing model, with vLLM inference support and a Hugging Face demo.
Fara1.5-27B is a browser computer-use agent that acts from screenshots by emitting clicks, typing, scrolling, URL visits, and web searches.
KAT-Coder-V2.5-Dev is a 35B-parameter MoE text-only model with 3B activated parameters.
GLM-5.2 adds a stable 1M-token context and multiple thinking effort levels for coding.
Solar Open 2 is a 250B-A15B MoE model that activates 15B parameters per token and targets tool calling and multi-step reasoning.
The Stack v3 is the largest open GitHub code dataset and includes full-repository context for pretraining code models.
This dataset contains 4,665 Pi-compatible trace sessions converted from Fable 5 coding-agent runs.
The corpus combines reasoning chains from DeepSeek, Qwen, and Gemma4-31B and keeps sequences within 5k tokens.
Nota AI released a 4-bit NVFP4 quantized version of Solar Open2 250B in llm-compressor format for vLLM, and the model requires a Blackwell GPU such as B200 or GB200.
Inflect-Micro-v2 is a sub-10M-parameter text-to-waveform speech model with fixed-voice English TTS, deterministic seeds, long-text handling, and CPU or CUDA inference.
This voice-agent pipeline combines VAD, STT, LLM, and TTS behind an OpenAI Realtime-compatible WebSocket API.
TRELLIS.2 is a 4B-parameter image-to-3D model that uses an O-Voxel sparse voxel structure for high-fidelity assets.
The package routes AI agents handling APKs, binaries, encrypted frontend JavaScript, CTFs, and pentesting targets to the right workflow.
gh-stack is a GitHub CLI extension for stacked branches and pull requests, and the repo also ships an AI agent skill for using it.
Kaneo is a project-management app built around fewer notifications, fewer buttons, and fewer workflows.
Invidious is an open-source YouTube front end with no ads, no tracking, and no JavaScript required.
Microsoft's curriculum offers 12 weeks and 24 lessons covering AI tools, quizzes, labs, and ethics.
Ansible is agentless infrastructure automation built around SSH and is used for configuration management, deployment, cloud provisioning, and multi-node orchestration.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.