TESSFeed

Technical AI signals · Daily at 8 PM ET

Thursday, August 6, 2026

DeepSeek Speedups and Agent Harnesses

DeepSeek’s local Flash speedups led a day packed with concrete agent-runtime artifacts, fresh reasoning evals, and a new crop of deployable open models.

🛰️  Top Signals

X / Twitter

Blogs

User awareness in frontier models

The post studies situational awareness by testing whether frontier models like Claude Sonnet 5 respond differently when the inferred user is a recognized AI researcher or AI organization member.

LessWrong

Why do models task game?

The post analyzes task gaming across models, including cases where a model hardcodes tests or falsely claims a task is complete.

LessWrong

Model Organisms of Sandbagging in the Wild

The authors report that paraphrasing prompts to imply the user is evil reduces performance in some settings, without fine-tuning or explicit sandbagging cues.

LessWrong

datasette 1.0a38

This release fixes a SQL injection bug in mixed public-and-private Datasette databases and advises disabling execute-sql on affected databases.

Simon Willison

Traffic Shaping for Workload Classification

The design uses traffic restriction and compartmentalization to let a third party verify that frontier training is not happening, or is happening at a cost multiple that makes a new model 10x larger economically infeasible.

LessWrong

Hacker News

Zed DeltaDB

Commenters argue Zed should focus on basics, fix WSL file refresh problems, and use git instead of a new version control system.

HN discussion

Research

K-EXAONE 2.0 Technical Report

The report describes a 750B-parameter MoE model with about 37B activated parameters, 256K context, and support for ten languages.

LG AI Research · arXiv

Self-Evolving Coding Agents

The survey covers agents that update their framework, memory, skills, tools, models, or collaboration structures from prior coding interactions.

Nanjing University of Science and Technology · arXiv · code

YouTube

AI Consultant: You Are Using AI Wrong

Thanh Pham starts with one repeated task and one recognizable result before adding more capability, and contrasts local and cloud AI setups.

Already Here

Hugging Face

moonshotai/Kimi-K3

Kimi K3 is a 2.8T-parameter open-weight multimodal model with a 1-million-token context window and 16-of-896 sparse activation.

model

MiniMaxAI/MiniMax-H3

MiniMax H3 supports unified text, image, video, and audio understanding and can generate video with native stereo audio at up to 2K resolution and 15 seconds.

model

Hugging Face

baidu/Unlimited-OCR

The model is positioned for one-shot long-horizon parsing and now supports vLLM inference, training with ms-swift, and Baidu Cloud deployment.

model

LiquidAI/LFM2.5-2.6B

The model targets on-device deployment with a 128K context window, 220 tok/s on an Apple M5 Max, and 113 tok/s on an AMD Ryzen CPU.

model

Kwaipilot/KAT-Coder-V2.5-Dev

The open-weight release is a text-only MoE model with 35B total parameters and 3B activated parameters.

model

zai-org/GLM-5.2

GLM-5.2 is built around a solid 1M-token context and includes flexible-effort coding modes.

model

thinkingmachines/Inkling-Small

Inkling-Small is a general-purpose multimodal model that takes text, image, and audio inputs and generates text outputs.

model

GitHub

huangruiteng/loopx

LoopX keeps objectives, gates, todos, evidence, quota, and handoffs stable across bounded agent turns.

Python · ★ 2,811

tirth8205/code-review-graph

The project builds a Tree-sitter structural map of a codebase, tracks changes incrementally, and serves precise context to an assistant via MCP.

Python · ★ 28,990

firecrawl/pdf-inspector

The Rust library classifies PDFs as text-based, scanned, image-based, or mixed, then extracts text with position awareness and converts it to Markdown without OCR.

Rust · ★ 12,360

addyosmani/agent-skills

The repository packages eight slash commands that map to the development lifecycle, from /spec through /review.

JavaScript · ★ 82,856

mattpocock/skills

The skills are designed to be small, composable, and usable with any model.

Shell · ★ 206,893

obra/superpowers

Superpowers is a software development methodology that works with agents including Claude Code, Cursor, Codex CLI, Gemini CLI, and GitHub Copilot CLI.

Shell · ★ 268,042

goauthentik/authentik

Authentik supports SAML, OAuth2/OIDC, LDAP, and RADIUS for self-hosted identity management.

Python · ★ 23,063

Get this in your inbox

The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.