Content

Frameworks

Design

Skills

Benchmarks

  • https://liveswebench.ai/
  • DeepSWE — Datacurve’s long-horizon coding-agent benchmark; 113 original tasks across 91 repos in 5 languages, run on mini-swe-agent (no PR/commit contamination)

Tools

  • claude-squad - Manage multiple AI terminal agents like Claude Code, Codex, and Amp
  • ml-intern - Hugging Face open-source ML engineer agent that reads papers, trains models, and ships ML models
  • cchistory — Claude Code version history viewer by Mario Zechner
  • deepsec — Vercel Labs’ agent-powered vulnerability scanner; regex matchers find candidate sites, then high-thinking models investigate and verify; supports distributed scanning (Apache-2.0)
  • Plannotator — local, open-source tool to review and annotate an agent’s plan and code before it runs, feeding inline comments back to the agent (Claude Code, Codex, Copilot…); VS Code + Obsidian integrations

Assorted