TESSFeed

Technical AI signals · Daily at 8 PM ET

Tuesday, July 14, 2026

Autoresearch Becomes a Product Surface

Today’s feed was dominated by agentic ML workflows turning into concrete products and benchmarks: NVIDIA pushed autoresearch from demo to playbook, Perplexity shipped wide research infrastructure, and model and safety claims kept colliding with real operator concerns.

🛰️  Top Signals

2

Perplexity opens up its research stack

Perplexity is exposing both the capability and the yardstick: Wide Research is now in the Agent API, and WANDR gives outsiders a benchmark for the deep-and-wide research behavior the product is optimized around.

@AravSrinivas

X / Twitter

Blogs

Open Distillation of Hereditary Traits

The post says distillation transferred traits such as Gemma 3 negative emotion into Qwen-base and Gemma 4 agentic misalignment into Nemotron Chat even after filtering prompts and rollouts that mentioned the trait.

Alignment Forum

Synthetic Scalable Oversight

The post proposes studying scalable oversight by training tiny models inside graphical abstractions of real-world problems and links code at AgoraForge.

LessWrong

lobste.rs is now running on SQLite

Simon Willison reports Lobsters moved from MariaDB to SQLite on a single VPS, with a 3.8GB primary database and lower CPU, memory, and hosting cost.

Simon Willison

simonw/pedalican

Simon Willison describes making a custom Codex Desktop pet by prompting GPT-5.6 Sol xhigh and using several rounds of gpt-image-2 to generate sprite assets.

Simon Willison

Hacker News

The Tower Keeps Rising

Commenters describe the essay as a critique of naive agent composability and compare its thesis to the Lisp Curse and Bipolar Lisp Programmer essays.

HN discussion

Research

Multi-Agent LLMs Fail to Explore Each Other

The paper says modern LLM agents show myopic, polarized interaction patterns in multi-agent settings and proposes Multi-Agent Contextual Exploration (MACE) to probe peers and reduce regret.

University of Wisconsin - Madison · arXiv

Evidence-Backed Video Question Answering

The paper defines E-VQA, requiring models to output both an answer and spatiotemporal evidence including temporal segments and dense tracked object masklets, and introduces the ST-Evidence benchmark.

Salesforce AI Research · arXiv

YouTube

Hugging Face

baidu/Unlimited-OCR

Unlimited-OCR targets one-shot long-horizon parsing, and the model card says it added vLLM inference support and has a Hugging Face Spaces demo.

model

OpenMOSS-Team/MOSS-Transcribe-Diarize

This 0.9B audio model does single-pass transcription, diarization, timestamps, and acoustic event awareness for recordings up to 90 minutes across 50+ languages.

model

nvidia/Nemotron-Labs-Audex-30B-A3B

Audex-30B-A3B extends a 30B MoE text model with discrete audio tokens and an audio encoder to handle speech recognition, translation, TTS, audio generation, and speech-to-speech.

model

Glint-Research/Fable-5-traces

This dataset releases 4,665 Pi-compatible coding-agent trace sessions converted from 60 source sessions, with 3,799 tool actions and a median 2,365 characters of chain-of-thought.

dataset

ByteDance-Seed/EdgeBench

EdgeBench evaluates agent learning over time on 134 real-world tasks, with 51 tasks and the full framework released publicly and 38,000+ hours of agent interaction analyzed.

dataset

netflix/Vera-Layered-Video-Dataset

This dataset supports Vera, a layered diffusion video editing framework that jointly generates an edit layer, alpha matte, and composite video to separate generated from preserved content.

dataset

smolagents/hf-realtime-voice

This Space is a minimal conversation app that swaps WebRTC SDP for a direct WebSocket connection to the Hugging Face speech-to-speech backend.

space

GitHub

Graphify-Labs/graphify

Graphify maps code, docs, PDFs, images, and videos into a queryable knowledge graph, and its code parsing is fully local via tree-sitter AST without an LLM.

Python · ★ 86,552

Shubhamsaboo/awesome-llm-apps

This repo packages 100+ open-source AI agents, agent skills, and RAG apps with end-to-end templates that work across Claude, Gemini, GPT, DeepSeek, Llama, and Qwen.

Python · ★ 120,989

mattpocock/skills

This repo offers small, composable agent skills intended for real engineering work and positioned as an alternative to more process-owning approaches like GSD, BMAD, and Spec-Kit.

Shell · ★ 170,559

Nutlope/hallmark

Hallmark is a design skill for Claude Code, Cursor, and Codex that applies one of 20 themes, four verbs, 57 slop-test gates, and a pre-emit self-critique.

CSS · ★ 6,266

HKUDS/Vibe-Trading

Vibe-Trading is a trading agent with API and MCP support, and the repo warns that a named X account, Virtuals project, and token contract are not official assets.

Python · ★ 22,972

virattt/ai-hedge-fund

This proof-of-concept uses multiple agents for trading decisions and is being rebuilt into a persistent system with backtestable alpha models and optional live execution.

Python · ★ 61,925

chenyme/grok2api

This Go project exposes OpenAI-style and Anthropic Messages-compatible APIs over pooled Grok Build, Grok Web, and Grok Console accounts with routing, quotas, and failover.

Go · ★ 5,892

Get this in your inbox

The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.