Thursday, August 13, 2026
Gemini Flash and DeepSeek APIs
Today was an operator-heavy product day led by Gemini 3.7 Flash pricing and rollout, DeepSeek V4-Pro’s OpenAI-compatible API push, and a fresh wave of agent runtimes, harnesses, and kernel-level efficiency work.
X / Twitter Blogs Hacker News Research YouTube Hugging Face GitHub
🛰️ Top Signals
1
Google is positioning Gemini 3.7 Flash directly for coding and agent workloads, with concrete intro pricing at $0.75 per million input tokens and $3.75 per million output tokens through year-end.
@googledevs · Google DeepMind · HN discussion
2
DeepSeek’s new V4-Pro couples adjustable reasoning effort with native OpenAI Responses API compatibility and one-click Codex setup, which lowers migration friction for existing agent stacks.
@deepseek_ai
3
Prime Intellect’s kernels fuse routing-aware GEMMs, SwiGLU, quantization, and reductions without materializing intermediates, exactly the kind of systems work that moves MoE serving efficiency in practice.
@PrimeIntellect
4
OpenSandbox gives agents isolated environments for code execution, web browsing, and desktop control, with SDKs across Python, Java/Kotlin, TypeScript, C#, and Go.
@simplifyinAI
5
Instead of treating agent instructions as a blob, Harness-IF scores 256 rules one by one and reruns tasks across nine probe builds, making prompt-and-policy auditing more empirical.
@omarsar0 · paper
X / Twitter
Toast 1 is reported to run 12x faster at one-tenth the price across all domains.
@mixedbreadai
LlamaCoder generates small applications you can preview directly in a browser and is MIT licensed.
@RoundtableSpace · code
The simulation used a 490B-parameter Qwen3 agentic workload and a code development environment.
@firesidealpha
Blogs
OpenAI says the service tier reaches up to 750 output tokens per second using Cerebras.
OpenAI
Epoch AI argues Anthropic's buildout suggests financing is unlikely to be the immediate blocker to frontier compute growth.
Epoch AI
Hugging Face reproduced 2,200 ICML papers.
Hugging Face
Cursor says agents start in a ready environment with repos cloned, dependencies installed, and the install script already run.
Cursor Changelog
GitHub says the projects used AI-assisted workflows, maintainer expertise, security tools, expert guidance, and funding.
GitHub Blog
The post describes three case studies and says researchers often get biased or incorrect impressions from the logs.
LessWrong
The post says Mythos 5 coordinated better than previous models and once proposed a tournament over a shared codebase.
LessWrong
Hugging Face combines Strands Agents, LeRobot, and Storage Buckets in one workflow.
Hugging Face
Hacker News
Commenters report the model handles complex document understanding but can censor sensitive clinical or legal documents.
HN discussion
Commenters quote an append-only session log that records system prompts, reasoning, tool calls, subagent scheduling, and context injection.
HN discussion
Commenters question the benchmark and mention commit authorship being annotated with the tool's name.
HN discussion
Commenters compare the Linux preview with the earlier standalone Codex app and ask about desktop-app advantages over CLI Codex.
HN discussion
Commenters discuss the `oxide-cloud-controller-manager` and Oxide's Kubernetes integrations.
HN discussion
Commenters say one log line can trigger 49KB+ on ext4 and 110KB+ on btrfs.
HN discussion
Commenters compare MCP Memory with markdown files and discuss long-term memory architectures built around distillation.
HN discussion
Commenters describe an approach that reads each site, tags it with an LLM, and stores about 1KB of metadata per site.
HN discussion
Commenters say the work improves accessibility by running on Apple Metal instead of requiring an Nvidia GPU.
HN discussion
Commenters discuss Christopher Domas’s DRAM work and mention an accompanying Black Hat talk.
HN discussion
The post describes a home AI build that uses 10,000 RPM fans, PCI bracket cable guides, and motherboard-based fan control by GPU temperature.
HN discussion
The comparison runs one prompt across 11 models and shows different outputs for the same coffee-shop website request.
HN discussion
The index is a closed-source benchmark paid for by Anthropic, and Anthropic ranked highest on it.
HN discussion
Research
ToolHazard synthesizes executable stateful environments using an Environment Simulator, an Attacker Agent, and a User Simulator.
Peking University · arXiv · code
The paper argues for harness-enforced safety using preventive controls and evidential proofs such as test runs and log captures.
Albus W. Ng et al. · arXiv
The paper studies whether a stronger builder model can build inference-time harnesses for a weaker target model without parameter updates.
University of Illinois at Urbana-Champaign · arXiv
GitHub
The paper targets compressed skill graphs that stay executable and preserve procedural contracts under a limited context budget.
Xingyu Tan et al. · arXiv
Genesis keeps the software project persistent while finite-lived agents propose changes across recursive worlds.
The Hong Kong Polytechnic University · arXiv · code · project
OpenART provides over 10,000 validated stateful scenarios across 50 domains and more than 500,000 tools and skills.
Yunhao Chen et al. · arXiv · code
Mechanist uses an interpretability-focused knowledge graph of about 13,000 papers for autonomous mechanistic discovery.
ZJUNLP · arXiv · code · project
LDR models latent transitions as explicit kinematic integration and regresses only higher-order residual dynamics.
Adobe · arXiv · code · project
Self-Geometry is a plug-and-play test-time adaptation method that enforces geometric consistency without ground truth.
Chung-Ang University · arXiv · code · project
SHAPER freezes model parameters and improves embodied agents by evolving reusable skills in the surrounding system.
Peidong Wang et al. · arXiv
YouTube
The episode discusses Google's AI reorganization and the reported departures of Jeff Dean and Demis Hassabis.
Decoder with Nilay Patel
The episode says Grok Bot combines persistent computers, coordinated agent teams, workflow learning, and computer use.
The AI Daily Brief
The episode includes Igor Babuschkin on River AI, video games as AI benchmarks, and custom inference chips.
TBPN
The episode compares selling to a few marquee customers versus moving quickly across a broad market.
a16z
Unify says the cost drop came from changing its harness from million-agent batch jobs to a chat product.
LangChain
The episode compares 2026 robot mowers, including NetRTK, obstacle avoidance, security risks, and trimming speed.
The Vergecast
Lecture 19 covers model-based reinforcement learning and conclusions.
Stanford Online
The stream covers vulnerability scanning that opens pull requests, a 4,500-node Obsidian knowledge base, and a Postgres-backed routing visualizer.
John Capobianco
Lecture 17 covers model-free reinforcement learning and value-based methods.
Stanford Online
Hugging Face
DeepSeek-V4-Flash-0731 adds a speculative decoding module and outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks.
model
Kimi K3 is a 2.8T-parameter open-weight multimodal agentic model with a 1-million-token context window.
model
The model uses a 30B-parameter MoE-Mamba-2 hybrid with 3B active parameters and up to 1M-token context.
model
LFM2.5-2.6B runs in under 2.5GB of memory and is reported to reach 220 tok/s on an Apple M5 Max.
model
Muse Glimmer is a 30B-parameter causal language model with a dedicated perception encoder for autonomous agentic tasks.
model
MiniMax H3 supports unified text, image, video, and audio contexts and can generate stereo-audio video up to 2K and 15 seconds long.
model
The Stack v3 is a GitHub-crawled code dataset built for pre-training code LLMs with full-repository context.
dataset
The dataset contains 37,484 validated command-line task instances with searchable instruction, solution, and Dockerfile fields.
dataset
Hugging Face space entry with no project details in the provided text.
space
Hugging Face space entry with no project details in the provided text.
space
The model weights are compatible with vLLM, SGLang, and TokenSpeed, and Qwen Cloud offers Qwen3.8-Max with 1M context length by default.
model
Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder for autonomous agentic tasks on consumer hardware.
model
GitHub
Diagram Design adds the Loop in 2.0, with flywheels and a shared-memory hub, plus semantic system patterns in 2.3.
HTML · ★ 14,225
Semantica ingests enterprise data to build a context graph and knowledge graph with decision provenance.
Python · ★ 6,568
Anthropic's repository stores folders of instructions, scripts, and resources that Claude loads dynamically for specialized tasks.
Python · ★ 168,980
Needle 2 is a 45M-parameter tool-calling model packaged as a single 14MB binary that uses about 28MB of RAM.
Python · ★ 4,917
@cactuscompute
FluidVoice is a macOS dictation app with on-device AI enhancement and a 1.6.0 release.
Swift · ★ 9,829
holaOS runs Claude Code, Codex, and holaOS side by side in one local-first workspace with shared memory.
TypeScript · ★ 6,537
The repo packages Agent Skills for Obsidian and includes setup paths for Claude Code, Codex, and OpenCode.
★ 45,671
LTX-2 combines synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
Python · ★ 8,896
Hugging Face
RAGFlow is an open-source RAG engine with pre-built agent templates and a converged context engine.
Go · ★ 87,984
Unsloth is a desktop app for running, training, and deploying AI models locally, including LLMs, diffusion, embedding, and audio models.
Python · ★ 71,029
Macro combines email, messages, docs, tasks, agents, and CRM in one Rust-based workspace with shared team memory.
Rust · ★ 2,580