Today centered on deployable model and agent infrastructure updates led by Gemini 3.8 Flash landing in tooling, Chrome DevTools over MCP, and a concrete new multi-day coding-agent loop in Harness-of-Harness.
Coverage update: X was unavailable for this edition.
Between the llm-gemini 0.34 update and the launch discussion, Gemini 3.8 Flash shows up as an immediately usable fast model with low/medium/high thinking levels and positive reports on web code generation cost and latency.
ChromeDevTools/chrome-devtools-mcp turns Chrome into an MCP surface for automation, debugging, and performance analysis, which is exactly the kind of concrete tool bridge agent operators can wire into existing loops now.
The HoH paper makes the agent pattern explicit: repeated planning, coding, and testing cycles with implementation-time tests separated from independent evaluation, giving teams a cleaner scaffold for long-running software agents.
Cursor’s new self-hosted machines keep tool execution inside your network, a practical shift for teams that want agent workflows without sending runtime actions outside their own perimeter.
NVIDIA’s co-design post is a concrete inference optimization artifact, showing how speculative decoding can speed LLM serving without changing output accuracy.
Model diff amplification, also called LDA, helps identify rare, unexpected behavior changes from a training run and can be used for red-teaming, evaluation, and monitoring.
Anthropic now publishes historical Claude consumer app system prompts, including model-specific pages such as Haiku 4.5 with October 15, 2025 and January 18, 2026 revisions.
The T-Tech paper consolidates traffic from more than 200 internal applications onto one model by closing gaps in instruction following, function calling, and internal task distribution.
DroneCATS puts an MLLM directly into a drone control loop and lets the model handle yaw, search, deliberation, and arrival without fine-tuning or function calling.
The paper separates execution protocols from task content so prompt edits do not corrupt routing, formatting, or termination signals in multi-agent pipelines.
Qwen-Drive-1.0 adds a 3D perception head and a Planning Expert for object detection, occupancy prediction, map segmentation, and future ego trajectories.
Moderna and Merck report positive Phase 3 results for an individualized melanoma treatment that encodes up to 34 tumor mutations into patient-specific mRNA.
Nikesh Arora says increasingly powerful AI models raise cybersecurity risks and that Palo Alto Networks uses small specialized models and AI-focused hiring.
Tom McGrath argues that interpretability can be used inside the training loop and discusses concept manifolds, reusable computation, and activation steering.
DeepSeek says this experimental multimodal model adds visual modules and improves multimodal agent capability while keeping text-only agent performance comparable to DeepSeek-V4-Flash-0731.
Anthropic released 1,440 de novo miniprotein binders against 16 targets, designed by two Claude models and characterized by two contract research organizations.
SIE serves more than 100 models from one API for search, retrieval, document-to-markdown conversion, structured output, content safety, and agent loops.
ECC says official installs come only from verified channels, including the GitHub repo, two npm packages, the GitHub App, the plugin slug, and the project website.