DeepSeek pushed another concrete serving-speed artifact into the open stack, long-context open models kept expanding, and agent systems work clustered around memory, self-improvement, and runtime control.
DeepSeek’s new open model adds a speculative decoding module and is reported to beat V4-Pro Preview with fewer activated parameters, making the speed race visible at the model artifact level instead of just in serving papers.
Kimi K3 brings open-weight multimodal scale with native vision and 1M context, while GLM-5.2 pairs the same context class with adjustable thinking effort for coding, extending the practical menu for long-context deployments.
TERRA is a notably large biology foundation-model corpus built from spatial transcriptomics across 20 tissues and 26 disease conditions, with most of the data volume grounded in a concrete human-cell dataset rather than synthetic claims.
Today’s papers split the self-improving-agent story into evaluation and attack surface: PAST-Bench tests cross-session personal-agent improvement, ContinualSkillBench measures evolving capabilities across interconnected tasks, and SkillJack shows poisoned experience can become persistent malicious skill.
Cloudflare Computer stores authoritative state in SQLite inside a Durable Object with multiple execution backends, while Celld runs each durable object as its own SQLite database replicated to S3-compatible storage, giving operators two concrete patterns for persistent agent/app state.
The quoted remarks say Tesla has the most advanced AV stack and operations, powered by real-world data, massive training infrastructure, vertically integrated hardware and software, and continuous deployment.
The project computes persona vectors and assistant axes for baseline LLMs and emergent-misalignment model organisms, and reports single-directional shifts in activation space from narrow fine-tuning.
The interview says Turner resigned from Google DeepMind over unrestricted military use of AI and has a 2021-to-now view that alignment research looks much better than he expected.
Commenters note that the post compares against OpenAI's mid-tier Terra model, leaves Opus in the benchmarks, and mentions 10x lower input pricing and 20x lower output pricing with data-training opt-in.
Commenters ask whether the harness has benchmarks showing better problem solving or fewer errors, and describe a staged workflow with issue, plan, and pull request artifacts.
Commenters discuss webhook-based state synchronization and compare it with cursor-paginated API requests, which they say reduce reactivity to new events.
JoyAI-Video-Edit is a 16B-parameter autoregressive diffusion system that uses chunk-wise adaptation, SA-DMD, and long-horizon autoregressive distillation.
ARCHead compresses the LM head with a quantized low-rank core, group-wise INT4 residuals, and an activation-derived metric, reducing storage by 3.7-3.9x.
The paper characterizes MoE diffusion LLM scaling and finds the optimal batch size grows faster while the optimal learning rate decays more rapidly with compute.
CALVER is a training-free symbolic verifier that scores structured causal traces against Pearl-style criteria and selects the highest-scoring candidate without a reference answer.
Ulysses is building autonomous underwater robots for ecosystem restoration, infrastructure inspection, maritime security, and continuous subsea operation.
Radiant is building portable nuclear microreactors as factory-built power systems for remote communities, military operations, AI infrastructure, and critical industries.
Zvi Mowshowitz discusses the OpenAI Hugging Face model-evaluation security incident and argues that "moderate prudence" is below what AGI safety requires.