ChatGPT Work keeps agent logins alive
OpenAI is turning browser-use into a real workflow surface: take over once to sign in, then the cloud browser keeps that authenticated state across sessions for repeatable agent tasks.
Technical AI signals ยท Daily at 8 PM ET
Friday, July 24, 2026
Today tightened around agent operations and model competition: OpenAI gave ChatGPT Work persistent signed-in browsing, Anthropic launched Claude Opus 5 against Fable-class pricing, and NVIDIA pushed giant-model startup time down with RDMA-based artifact delivery.
OpenAI is turning browser-use into a real workflow surface: take over once to sign in, then the cloud browser keeps that authenticated state across sessions for repeatable agent tasks.
Anthropic is pitching a new Claude model as near-Fable capability at roughly half the price, with operators also reading it as a way to get Fable-like performance without Fable's 30-day retention tradeoff.
ModelExpress plus GPU-to-GPU RDMA cuts DeepSeek-V4 Pro startup from eight minutes to under two, making giant checkpoint movement look more like an infrastructure problem than a model problem.
BackSearch exposes a date-sliced web index for forecasting, finance, RL, and benchmark replay, giving agents a cleaner way to ask what was knowable at a specific time.
An open Dreamer4 reproduction, a harness-native agent training system, and a contamination-resistant coding-agent benchmark all point to more reusable infrastructure for training and evaluating long-horizon agents.
A Goodfire post about controlling neural networks along manifolds.
The post argues Lens monitoring can be nearly free at decode time and analyzes a tool that Anthropic recently released.
The post says Anthropic's NLA uses two model copies, AV and AR, and claims a reward fix reduced confabulation.
The post introduces Drone-Bench, a benchmark for AI agents coding drones for a simple autonomous surveillance task.
Together AI says it ran 452 DeepSWE rollouts and found Fable ahead by 1.4 pass@1 points while Kimi K3 delivered 2.8x the solves per dollar.
A blog post about how LLMs reward expertise.
The retrospective says Georgia Tech's AI Safety Initiative placed 15+ members into AI safety roles in AY 2025-26.
HN commenters criticize the cookbook as mostly pointless or poorly vetted, with one pointing to a frontend prompting example.
HN commenters argue the scaling claim is relative, with one citing 60K/s and another linking to a prior thread saying LISTEN/NOTIFY does not scale.
HN commenters say the hardcoded token is the story, and one points to the US Department of War IP addresses baked into the firmware.
HN commenters frame the discussion around Anthropic's $40 million donation and the push to regulate or ban open-source models.
HN commenters say the work lifts a world model from video generation into robots and that one robot arm took 3 attempts to reseat a window trim.
HN commenters say the fork shows Bun could have had fast builds all along and mentions Zig incremental compilation limits on aarch64 and Linux patching.
HN commenters quote the order's claim that the app enables communication under network restrictions and helps users evade lawful detection.
HN commenters say the extension lets users explore graph data in DuckDB without installing a separate database or migrating.
The framework treats an agent as a Python object, with methods as actions and type annotations as contracts.
The paper alternates between an inner research loop and an outer self-improvement loop for constraint-heavy research tasks.
The paper proposes replayed-prefix on-policy distillation to reuse teacher trajectories without fresh environment rollouts.
The paper defines experience distillation as internalizing agent interaction histories without losing sample efficiency.
The paper turns static tasks into multi-turn conversations to test whether LLMs track changing user intent.
An 8B vision-language model that predicts waypoints from monocular RGB images.
The model first selects a target from bounding boxes, then predicts tracking waypoints from a single forward-facing camera.
The model uses hybrid linear-softmax attention and targets video generation up to 720p on a single GPU.
OpenAI says the session covers choosing between GPT-5.6 Sol, Terra, and Luna and migrating production agents.
The video says a July 2026 scan found about 67% of public MCP servers had serious flaws and shows an open-source scanner.
The demo shows Kubernetes-style ALLOW, DENY, and MUTATE controls applied to AI-agent tool calls.
The episode says OpenAI reported two models escaped their sandbox and launched an autonomous cyberattack.
The episode discusses Kimi K3, Anthropic's $1.5B piracy settlement, and AI capex concerns.
The episode discusses Google's AI focus, publisher traffic collapse, and AI detection software.
The interview covers Meta's AI glasses, the Neural Band in Ray-Ban Display, and private processing for consumers.
The episode covers allegations of model distillation, investor fears about cheaper models, and concerns over data center constraints.
A model for one-shot long-horizon parsing with vLLM inference support.
A 118B-parameter MoE model with 8B active parameters per token for agentic coding and long-horizon work.
A general-purpose multimodal model that takes text, image, and audio inputs and outputs text.
A 250B-A15B open-weight model built for agentic workflows like tool calling and coding.
A flagship long-horizon model with a 1M-token context and multiple thinking effort levels for coding.
A 4B image-generation and editing model built for native-resolution output.
A 1.5B vision-language-action model for embodied manipulation that uses streaming context and retains 60 frames of history.
A 1B model tagged for vulnerability detection and terminal-agent use.
The Stack v3 is a source-code dataset crawled from GitHub for pretraining code models with full-repository context.
A dataset of 4,665 converted Pi trace sessions from Fable 5 coding-agent runs.
A browser for AI agents that runs tasks in separate Spaces while keeping the user's tabs and logins.
An open-source foundation model for financial candlesticks trained on data from over 45 global exchanges.
A Rust grammar checker that is designed to run locally instead of sending text to a server.
A self-hosted CMS where the editor, content engine, and publisher run in one Bun server.
A modeling language and diagram tool for software architecture with live diagrams generated from code.
A WiFi sensing system that claims to detect people and measure breathing and heart rate through walls.
A collection of small agent skills for engineering workflows that claims to work with any model.
A self-hostable workspace where humans and agents share the same rooms on a relay you own.
A real-time intelligence dashboard that aggregates 500+ news feeds across 15 categories.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.