Data scientist
Working with data, notebooks, and models. You pair a coding agent with evaluation for when a model has to be right. 38 tools.
The shape of a stack that fits this role. Open one for the full diagram, the tools that fill it, and a worked example.
Write code with AI
Full stack & case study →I want an AI that writes and edits code with me.
flowchart LR
code(("Your code"))
agent["Coding agent"]
model["The model"]
herd["Run several at once"]
agent -->|reads and edits| code
agent -->|sends prompts to| model
herd -->|runs| agent Test and score my agents
Full stack & case study →I want to test how well my AI behaves before I ship it.
flowchart LR
app(("Your agent"))
eval["Evaluation"]
app -->|is tested by| eval Aider
AI pair programming in your terminal, git-native.
Augment Code
A context engine for large enterprise codebases.
Claude Code
Anthropic's agentic CLI coding assistant.
Cline
Autonomous coding agent inside VS Code, with approval gates.
Continue
Open-source AI code assistant for VS Code and JetBrains.
Cursor
The AI-first code editor built on VS Code.
Devin
Cognition's autonomous software engineer.
Gemini CLI
Google's open-source terminal agent for Gemini.
GitHub Copilot
GitHub's AI pair programmer across editor and CLI.
Goose
Block's open-source local AI agent.
Hermes Agent
The self-improving AI agent built by Nous Research.
Kilo Code
Open-source coding agent consolidating Cline and Roo.
Open Interpreter
Let an LLM run code on your machine.
OpenAI Codex
OpenAI's open-source agentic coding CLI.
OpenCode
A model-agnostic, local-first coding agent TUI.
OpenCode Autolearn
Self-improvement engine for OpenCode agents.
OpenHands
Open-source autonomous software engineer (formerly OpenDevin).
Pi
A minimal, extensible terminal coding-agent harness.
Replit Agent
Build and ship apps conversationally on Replit.
Roo Code
A community fork of Cline with extra modes.
Tabby
Self-hosted AI coding assistant for teams.
Void
An open-source Cursor alternative.
Windsurf
Codeium's AI-native IDE (Cascade agent).
Zed (AI)
The high-performance editor with built-in AI.
AgentOps
Observability built specifically for AI agents.
Arize Phoenix
Open-source LLM tracing and eval.
Comet Opik
Open-source eval and tracing from Comet.
Laminar
Open-source observability and eval for AI agents.
Langfuse
The most-used open-source LLM observability tool.
LangSmith
LangChain's observability and eval platform.
Weights and Biases Weave
LLM tracing and eval inside W&B.
Braintrust
Evals and prompt playground for serious teams.
DeepEval (Confident AI)
Open-source unit tests for LLMs.
Galileo
Guardrails and evaluation for production LLMs.
Latitude
Open-source prompt management and eval.
Maxim AI
Evaluation and simulation for AI agents.
Patronus AI
Automated evaluation and guardrails for LLMs.
promptfoo
Test and red-team LLM apps, prompts, and agents.