S05. TodoWrite — An Agent Without a Plan Drifts Off Course
S05. TodoWrite — An Agent Without a Plan Drifts Off Course
The Problem
Give the Agent a complex task: "Rename all Python files to snake_case, run tests, and fix failures."
The Agent starts working, renames 3 files, runs a test, finds 2 failures, starts fixing. While fixing, it forgets the original goal was "rename to snake_case", the test failures have consumed all its attention.
The longer the conversation, the worse it gets: tool results keep filling the context, diluting the system prompt's influence. A 10-step refactoring: after steps 1-3, the Agent starts improvising because steps 4-10 have been pushed out of its attention.
The Solution
S05 keeps the tool dispatch, permissions, and hooks from S04, then adds todo_write and a reminder counter. todo_write only updates planning state; the existing tools still perform the work.
The new tool uses the same TOOL_HANDLERS[block.name] dispatch path. After three consecutive tool-use rounds without todo_write, the harness adds a reminder to that round's tool results.
How It Works
TodoManager owns the in-memory list, validates updates, and renders the state returned to the model. run_todo_write also prints that state in the terminal:
class TodoManager:
def __init__(self):
self.items = []
def update(self, todos: list | str) -> str:
# Parse and validate before replacing the current list.
validated = []
...
self.items = validated
return self.render()
def render(self) -> str:
# [ ] pending, [>] in progress, [x] completed
...
TODO = TodoManager()
def run_todo_write(todos: list | str) -> str:
output = TODO.update(todos)
print(output)
return outputAn update may contain at most 20 items, each item needs non-empty content, and only one item may be in_progress. The string input path accepts JSON or a Python list representation without using eval.
The tool definition joins the other 5 in the dispatch map:
TOOLS = [
{"name": "bash", ...},
{"name": "read_file", ...},
{"name": "write_file", ...},
{"name": "edit_file", ...},
{"name": "glob", ...},
# s05: new entry
{"name": "todo_write", "description": "Create and manage a task list ...",
"input_schema": {
"type": "object",
"properties": {
"todos": {
"type": "array",
"items": {
"type": "object",
"properties": {
"content": {"type": "string"},
"status": {"type": "string", "enum": ["pending", "in_progress", "completed"]},
},
},
},
},
},
},
]
TOOL_HANDLERS["todo_write"] = run_todo_writeReminder: after three tool-use rounds without todo_write, the reminder is appended to the third round's results and the counter resets:
rounds_since_todo = 0 if used_todo else rounds_since_todo + 1
if rounds_since_todo >= 3:
results.append({
"type": "text",
"text": "<reminder>Update your todos.</reminder>",
})
rounds_since_todo = 0Typical flow when the Agent receives a task: first call todo_write to list all steps (all pending) → pick one step, set it to in_progress → complete it, set to completed → look at the next pending → continue.
Key insight: todo_write doesn't give the Agent any additional execution capability. What it adds is planning capability.
Changes from s04
| Component | Before (s04) | After (s05) | |-----------|-------------|-------------| | Tool count | 5 (bash, read, write, edit, glob) | 6 (+todo_write) | | Planning | None | Stateful TODO list + reminder | | SYSTEM prompt | Generic prompt | Added "plan before executing" guidance | | Loop | Tool dispatch and hooks | Same dispatch path, plus rounds_since_todo and reminder injection |
Try It
cd learn-claude-code
python s05_todo_write/code.pyTry these prompts:
Refactor s05_todo_write/example/hello.py: add type hints, docstrings, and a main guard(should list 3 steps first, then execute)Create a Python package under s05_todo_write/example/demo_pkg with init.py, utils.py, and tests/test_utils.pyReview Python files under s05_todo_write/example and fix any style issues
What to watch for: Was the first tool call todo_write? How many TODO steps were listed? Did statuses move from pending to in_progress / completed during execution?
What's Next
The Agent can plan now. But if a task is too large, say "refactor the entire auth module", a TODO list alone isn't enough. That task is itself a collection of dozens of subtasks that would drown in a single conversation's context.
→ s06 Subagent: Break large tasks into subtasks, each handled by an independent Agent with its own clean context, no cross-contamination.
S05 — Complete teaching code
#!/usr/bin/env python3
"""
s05_todo_write.py - TodoWrite
The model tracks its progress through a TodoManager. After three rounds
without an update, the harness adds a reminder alongside the tool results.
+----------+ +-------+ +--------------+
| User | ---> | LLM | ---> | Tools |
| prompt | | | | + todo_write |
+----------+ +---^---+ +------+-------+
| | update
| +------v----------+
| | TodoManager |
| | [ ] pending |
| | [>] in progress |
| | [x] completed |
| +------+----------+
| tool_result |
+-----------------+
rounds_since_todo >= 3 -> add <reminder>
"""
import ast
import json
import os
import subprocess
from pathlib import Path
try:
import readline
readline.parse_and_bind('set bind-tty-special-chars off')
except ImportError:
pass
from anthropic import Anthropic
from dotenv import load_dotenv
load_dotenv(override=True)
if os.getenv("ANTHROPIC_BASE_URL"):
os.environ.pop("ANTHROPIC_AUTH_TOKEN", None)
WORKDIR = Path.cwd()
client = Anthropic(base_url=os.getenv("ANTHROPIC_BASE_URL"))
MODEL = os.environ["MODEL_ID"]
# s05 change: SYSTEM prompt adds planning guidance
SYSTEM = (
f"You are a coding agent at {WORKDIR}. "
"Before starting any multi-step task, use todo_write to plan your steps. "
"Update status as you go."
)
# -- Tool implementations from s02-s04 --
def run_bash(command: str) -> str:
try:
r = subprocess.run(command, shell=True, cwd=WORKDIR,
capture_output=True, text=True, timeout=120)
out = (r.stdout + r.stderr).strip()
return out[:50000] if out else "(no output)"
except subprocess.TimeoutExpired:
return "Error: Timeout (120s)"
def run_read(path: str, limit: int | None = None) -> str:
try:
lines = (WORKDIR / path).resolve().read_text().splitlines()
if limit and limit < len(lines):
lines = lines[:limit] + [f"... ({len(lines) - limit} more lines)"]
return "\n".join(lines)
except Exception as e:
return f"Error: {e}"
def run_write(path: str, content: str) -> str:
try:
file_path = (WORKDIR / path).resolve()
file_path.parent.mkdir(parents=True, exist_ok=True)
file_path.write_text(content)
return f"Wrote {len(content)} bytes to {path}"
except Exception as e:
return f"Error: {e}"
def run_edit(path: str, old_text: str, new_text: str) -> str:
try:
file_path = (WORKDIR / path).resolve()
text = file_path.read_text()
if old_text not in text:
return f"Error: text not found in {path}"
file_path.write_text(text.replace(old_text, new_text, 1))
return f"Edited {path}"
except Exception as e:
return f"Error: {e}"
def run_glob(pattern: str) -> str:
import glob as g
try:
results = []
for match in g.glob(pattern, root_dir=WORKDIR):
if (WORKDIR / match).resolve().is_relative_to(WORKDIR):
results.append(match)
return "\n".join(results) if results else "(no matches)"
except Exception as e:
return f"Error: {e}"
# -- New in s05: structured state the model updates --
class TodoManager:
def __init__(self):
self.items: list[dict] = []
def update(self, todos: list | str) -> str:
if isinstance(todos, str):
try:
todos = json.loads(todos)
except json.JSONDecodeError:
try:
todos = ast.literal_eval(todos)
except (SyntaxError, ValueError) as e:
raise ValueError("todos must be a list or JSON array string") from e
if not isinstance(todos, list):
raise ValueError("todos must be a list")
if len(todos) > 20:
raise ValueError("Max 20 todos allowed")
validated = []
in_progress_count = 0
for index, todo in enumerate(todos):
if not isinstance(todo, dict):
raise ValueError(f"todos[{index}] must be an object")
content = str(todo.get("content", "")).strip()
status = str(todo.get("status", "pending")).lower()
if not content:
raise ValueError(f"todos[{index}] requires content")
if status not in ("pending", "in_progress", "completed"):
raise ValueError(f"todos[{index}] has invalid status '{status}'")
if status == "in_progress":
in_progress_count += 1
validated.append({"content": content, "status": status})
if in_progress_count > 1:
raise ValueError("Only one todo can be in_progress at a time")
self.items = validated
return self.render()
def render(self) -> str:
if not self.items:
return "No todos."
lines = []
for todo in self.items:
marker = {
"pending": "[ ]",
"in_progress": "[>]",
"completed": "[x]",
}[todo["status"]]
lines.append(f"{marker} {todo['content']}")
done = sum(todo["status"] == "completed" for todo in self.items)
lines.append(f"\n({done}/{len(self.items)} completed)")
return "\n".join(lines)
TODO = TodoManager()
def run_todo_write(todos: list | str) -> str:
try:
output = TODO.update(todos)
except ValueError as e:
return f"Error: {e}"
print(f"\n\033[33m## Current Tasks\033[0m\n{output}")
return output
TOOLS = [
{"name": "bash", "description": "Run a shell command.",
"input_schema": {"type": "object", "properties": {"command": {"type": "string"}}, "required": ["command"]}},
{"name": "read_file", "description": "Read file contents.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"}, "limit": {"type": "integer"}}, "required": ["path"]}},
{"name": "write_file", "description": "Write content to a file.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"}, "content": {"type": "string"}}, "required": ["path", "content"]}},
{"name": "edit_file", "description": "Replace exact text in a file once.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"}, "old_text": {"type": "string"}, "new_text": {"type": "string"}}, "required": ["path", "old_text", "new_text"]}},
{"name": "glob", "description": "Find files matching a glob pattern.",
"input_schema": {"type": "object", "properties": {"pattern": {"type": "string"}}, "required": ["pattern"]}},
# s05: new tool
{"name": "todo_write", "description": "Create and manage a task list for your current coding session.",
"input_schema": {"type": "object", "properties": {"todos": {"type": "array", "maxItems": 20, "items": {"type": "object", "properties": {"content": {"type": "string", "minLength": 1}, "status": {"type": "string", "enum": ["pending", "in_progress", "completed"]}}, "required": ["content", "status"]}}}, "required": ["todos"]}},
]
TOOL_HANDLERS = {
"bash": run_bash, "read_file": run_read, "write_file": run_write,
"edit_file": run_edit, "glob": run_glob, "todo_write": run_todo_write,
}
# -- Hook system from s04 --
HOOKS = {"UserPromptSubmit": [], "PreToolUse": [], "PostToolUse": [], "Stop": []}
def register_hook(event: str, callback):
HOOKS[event].append(callback)
def trigger_hooks(event: str, *args):
for callback in HOOKS[event]:
result = callback(*args)
if result is not None:
return result
return None
DENY_LIST = ["rm -rf /", "sudo", "shutdown", "reboot", "mkfs", "dd if="]
DESTRUCTIVE = ["rm ", "> /etc/", "chmod 777"]
def permission_hook(block):
"""PreToolUse: s03 permission logic, registered as an s04 hook."""
if block.name == "bash":
command = block.input.get("command", "")
for pattern in DENY_LIST:
if pattern in command:
print(f"\n\033[31m[blocked] '{pattern}'\033[0m")
return "Permission denied by deny list"
for keyword in DESTRUCTIVE:
if keyword in command:
print(f"\n\033[33m[permission] Potentially destructive command\033[0m")
print(f" Tool: {block.name}({block.input})")
choice = input(" Allow? [y/N] ").strip().lower()
if choice not in ("y", "yes"):
return "Permission denied by user"
if block.name in ("read_file", "write_file", "edit_file"):
path = block.input.get("path", "")
if not (WORKDIR / path).resolve().is_relative_to(WORKDIR):
print(f"\n\033[33m[permission] Access outside workspace\033[0m")
print(f" Tool: {block.name}({block.input})")
choice = input(" Allow? [y/N] ").strip().lower()
if choice not in ("y", "yes"):
return "Permission denied by user"
return None
def log_hook(block):
"""PreToolUse: log every tool call."""
args_preview = str(list(block.input.values())[:2])[:60]
print(f"\033[90m[HOOK] {block.name}({args_preview})\033[0m")
return None
def large_output_hook(block, output):
"""PostToolUse: warn on large output."""
if len(str(output)) > 100000:
print(f"\033[33m[HOOK] Large output from {block.name}: {len(str(output))} chars\033[0m")
return None
def context_inject_hook(query: str):
"""UserPromptSubmit: log working directory."""
print(f"\033[90m[HOOK] UserPromptSubmit: working in {WORKDIR}\033[0m")
return None
def summary_hook(messages: list):
"""Stop: print tool call count."""
tool_count = sum(1 for m in messages
for b in (m.get("content") if isinstance(m.get("content"), list) else [])
if isinstance(b, dict) and b.get("type") == "tool_result")
print(f"\033[90m[HOOK] Stop: session used {tool_count} tool calls\033[0m")
return None
register_hook("UserPromptSubmit", context_inject_hook)
register_hook("PreToolUse", permission_hook)
register_hook("PreToolUse", log_hook)
register_hook("PostToolUse", large_output_hook)
register_hook("Stop", summary_hook)
# -- Agent loop with the reminder counter --
def agent_loop(messages: list):
rounds_since_todo = 0
while True:
response = client.messages.create(
model=MODEL, system=SYSTEM, messages=messages,
tools=TOOLS, max_tokens=8000,
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
force = trigger_hooks("Stop", messages)
if force:
messages.append({"role": "user", "content": force})
continue
return
results = []
used_todo = False
for block in response.content:
if block.type != "tool_use":
continue
blocked = trigger_hooks("PreToolUse", block)
if blocked:
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": str(blocked)})
continue
handler = TOOL_HANDLERS.get(block.name)
try:
output = handler(**block.input) if handler else f"Unknown: {block.name}"
except Exception as e:
output = f"Error: {e}"
trigger_hooks("PostToolUse", block, output)
if block.name == "todo_write":
used_todo = True
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": str(output)})
rounds_since_todo = 0 if used_todo else rounds_since_todo + 1
if rounds_since_todo >= 3:
results.append({"type": "text",
"text": "<reminder>Update your todos.</reminder>"})
rounds_since_todo = 0
messages.append({"role": "user", "content": results})
if __name__ == "__main__":
print("s05: TodoWrite - plan before execution")
print("Enter a question, press Enter to send. Type q to quit.\n")
history = []
while True:
try:
query = input("\033[36ms05 >> \033[0m")
except (EOFError, KeyboardInterrupt):
break
if query.strip().lower() in ("q", "exit", ""):
break
trigger_hooks("UserPromptSubmit", query)
history.append({"role": "user", "content": query})
agent_loop(history)
for block in history[-1]["content"]:
if getattr(block, "type", None) == "text":
print(block.text)
print()
Try it — TodoWrite scenario
A planning tool makes the agent's intended work visible and keeps long tasks from drifting.