skills/autonomous-agents/SKILL.md
Autonomous agents are AI systems that can independently decompose goals, plan actions, execute tools, and self-correct without constant human guidance. The challenge isn't making them capable - it's making them reliable. Every extra decision multiplies failure probability.
npx skillsauth add Regtransfers/agency-agents-mcp autonomous-agentsInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
@ Autonomous Agents
Autonomous agents are AI systems that can independently decompose goals, plan actions, execute tools, and self-correct without constant human guidance. The challenge isn't making them capable - it's making them reliable. Every extra decision multiplies failure probability.
This skill covers agent loops (ReAct, Plan-Execute), goal decomposition, reflection patterns, and production reliability. Key insight: compounding error rates kill autonomous agents. A 95% success rate per step drops to 60% by step 10. Build for reliability first, autonomy second.
2025 lesson: The winners are constrained, domain-specific agents with clear boundaries, not "autonomous everything." Treat AI outputs as proposals, not truth.
@ Principles
@ Capabilities
@ Scope
@ Tooling
@ Frameworks
@ Patterns
@ Patterns
@ ReAct Agent Loop
Alternating reasoning and action steps
When to use: Interactive problem-solving, tool use, exploration
@ REACT PATTERN:
""" The ReAct loop:
Key: Explicit reasoning traces make debugging possible """
@ Basic ReAct Implementation """ from langchain.agents import createreactagent from langchain_openai import ChatOpenAI
@ Define the ReAct prompt template react_prompt = ''' Answer the question using the following format:
Question: the input question Thought: reason about what to do Action: tool_name Action Input: input to the tool Observation: result of the action ... (repeat Thought/Action/Observation as needed) Thought: I now know the final answer Final Answer: the answer '''
@ Create the agent agent = createreactagent( llm=ChatOpenAI(model="gpt-4o"), tools=tools, prompt=react_prompt, )
@ Execute with step limit result = agent.invoke( {"input": query}, config={"max_iterations": 10} # Prevent runaway loops ) """
@ LangGraph ReAct (Production) """ from langgraph.prebuilt import createreactagent from langgraph.checkpoint.postgres import PostgresSaver
@ Production checkpointer checkpointer = PostgresSaver.fromconnstring( os.environ["POSTGRES_URL"] )
agent = createreactagent( model=llm, tools=tools, checkpointer=checkpointer, # Durable state )
@ Invoke with thread for state persistence config = {"configurable": {"thread_id": "user-123"}} result = agent.invoke({"messages": [query]}, config) """
@ Plan-Execute Pattern
Separate planning phase from execution
When to use: Complex multi-step tasks, when full plan visibility matters
@ PLAN-EXECUTE PATTERN:
""" Two-phase approach:
Advantages:
Disadvantages:
@ LangGraph Plan-Execute """ from langgraph.prebuilt import createplanandexecuteagent
@ Planner creates the task list planner_prompt = ''' For the given objective, create a step-by-step plan. Each step should be atomic and actionable. Format: numbered list of steps. '''
@ Executor handles individual steps executor_prompt = ''' You are executing step {step_number} of the plan. Previous results: {previous_results} Current step: {current_step} Execute this step using available tools. '''
agent = createplanandexecuteagent( planner=planner_llm, executor=executor_llm, tools=tools, replanonerror=True, # Re-plan if step fails )
@ Human approval of plan config = { "configurable": { "thread_id": "task-456", }, "interrupt_before": ["execute"], # Pause before execution }
@ First call creates plan plan = agent.invoke({"objective": goal}, config)
@ Review plan, then continue if human_approves(plan): result = agent.invoke(None, config) # Continue from checkpoint """
@ Decomposition Strategies """ @ Decomposition-First: Plan everything, then execute @ Best for: Stable tasks, need full plan approval
@ Interleaved: Plan one step, execute, repeat @ Best for: Dynamic tasks, learning as you go
def interleavedexecute(goal, maxsteps=10): state = {"goal": goal, "completed": [], "remaining": [goal]}
for step in range(max_steps): @ Plan next action based on current state nextaction = planner.plannext(state)
if next_action == "DONE": break
@ Execute and update state result = executor.execute(next_action) state["completed"].append((next_action, result))
@ Re-evaluate remaining work state["remaining"] = planner.reassess(state)
return state """
@ Reflection Pattern
Self-evaluation and iterative improvement
When to use: Quality matters, complex outputs, creative tasks
@ REFLECTION PATTERN:
""" Self-correction loop:
Also called: Evaluator-Optimizer, Self-Critique """
@ Basic Reflection """ def reflectandimprove(task, max_iterations=3): @ Initial generation output = generator.generate(task)
for i in range(max_iterations): @ Evaluate output critique = evaluator.critique( task=task, output=output, criteria=[ "Correctness", "Completeness", "Clarity", ] )
if critique["passes_all"]: return output
@ Refine based on critique output = generator.refine( task=task, previous_output=output, critique=critique["feedback"], )
return output # Best effort after max iterations """
@ LangGraph Reflection """ from langgraph.graph import StateGraph
def buildreflectiongraph(): graph = StateGraph(ReflectionState)
@ Nodes graph.addnode("generate", generatenode) graph.addnode("reflect", reflectnode) graph.addnode("output", outputnode)
@ Edges graph.add_edge("generate", "reflect") graph.addconditionaledges( "reflect", should_continue, { "continue": "generate", # Loop back "end": "output", } )
return graph.compile()
def should_continue(state): if state["iteration"] >= 3: return "end" if state["score"] >= 0.9: return "end" return "continue" """
@ Separate Evaluator (More Robust) """ @ Use different model for evaluation to avoid self-bias generator = ChatOpenAI(model="gpt-4o") evaluator = ChatOpenAI(model="gpt-4o-mini") # Different perspective
@ Or use specialized evaluators from langchain.evaluation import load_evaluator evaluator = load_evaluator("criteria", criteria="correctness") """
@ Guardrailed Autonomy
Constrained agents with safety boundaries
When to use: Production systems, critical operations
@ GUARDRAILED AUTONOMY:
""" Production agents need multiple safety layers:
@ Multi-Layer Guardrails """ class GuardedAgent: def init(self, agent, config): self.agent = agent self.maxcost = config.get("maxcost_usd", 1.0) self.maxsteps = config.get("maxsteps", 10) self.allowedactions = config.get("allowedactions", []) self.requireapproval = config.get("requireapproval", [])
async def execute(self, goal): total_cost = 0 steps = 0
while steps < self.max_steps: @ Get next action action = await self.agent.plan_next(goal)
@ Validate action is allowed if action.name not in self.allowed_actions: raise ActionNotAllowedError(action.name)
@ Check if approval needed if action.name in self.require_approval: approved = await self.requesthumanapproval(action) if not approved: return {"status": "rejected", "action": action}
@ Estimate cost estimatedcost = self.estimatecost(action) if totalcost + estimatedcost > self.max_cost: raise CostLimitExceededError(total_cost)
@ Execute with rollback capability checkpoint = await self.save_checkpoint() try: result = await self.agent.execute(action) totalcost += self.actualcost(action) steps += 1 except Exception as e: await self.rollback_to(checkpoint) raise
if result.is_complete: break
return {"status": "complete", "totalcost": totalcost} """
@ Least Privilege Principle """ @ Define minimal permissions per task type TASK_PERMISSIONS = { "research": ["websearch", "readfile"], "coding": ["readfile", "writefile", "run_tests"], "admin": ["all"], # Rarely grant this }
def createscopedagent(task_type): allowed = TASKPERMISSIONS.get(tasktype, []) tools = [t for t in ALL_TOOLS if t.name in allowed] return Agent(tools=tools) """
@ Cost Control """ @ Context length grows quadratically in cost @ Double context = 4x cost
def trimcontext(messages, maxtokens=4000): @ Keep system message and recent messages system = messages[0] recent = messages[-10:]
@ Summarize middle if needed if len(messages) > 11: middle = messages[1:-10] summary = summarize(middle) return [system, summary] + recent
return messages """
@ Durable Execution Pattern
Agents that survive failures and resume
When to use: Long-running tasks, production systems, multi-day processes
@ DURABLE EXECUTION:
""" Production agents must:
LangGraph 1.0 provides this natively. """
@ LangGraph Checkpointing """ from langgraph.checkpoint.postgres import PostgresSaver from langgraph.graph import StateGraph
@ Production checkpointer (not MemorySaver!) checkpointer = PostgresSaver.fromconnstring( os.environ["POSTGRES_URL"] )
@ Build graph with checkpointing graph = StateGraph(AgentState) @... add nodes and edges...
agent = graph.compile(checkpointer=checkpointer)
@ Each invocation saves state config = {"configurable": {"thread_id": "long-task-789"}}
@ Start task agent.invoke({"goal": complex_goal}, config)
@ If server dies, resume later: state = agent.get_state(config) if not state.is_complete: agent.invoke(None, config) # Continues from checkpoint """
@ Human-in-the-Loop Interrupts """ @ Pause at specific nodes agent = graph.compile( checkpointer=checkpointer, interruptbefore=["criticalaction"], # Pause before interrupt_after=["validation"], # Pause after )
@ First invocation pauses at interrupt result = agent.invoke({"goal": goal}, config)
@ Human reviews state state = agent.get_state(config) if human_approves(state): @ Continue from pause point agent.invoke(None, config) else: @ Modify state and continue agent.update_state(config, {"approved": False}) agent.invoke(None, config) """
@ Time-Travel Debugging """ @ LangGraph stores full history history = list(agent.getstatehistory(config))
@ Go back to any previous state past_state = history[5] agent.updatestate(config, paststate.values)
@ Replay from that point with modifications agent.invoke(None, config) """
@ Sharp Edges
@ Error Probability Compounds Exponentially
Severity: CRITICAL
Situation: Building multi-step autonomous agents
Symptoms: Agent works in demos but fails in production. Simple tasks succeed, complex tasks fail mysteriously. Success rate drops dramatically as task complexity increases. Users lose trust.
Why this breaks: Each step has independent failure probability. A 95% success rate per step sounds great until you realize:
This is the fundamental limit of autonomous agents. Every additional step multiplies failure probability.
Recommended fix:
@ Reduce step count @ Combine steps where possible @ Prefer fewer, more capable steps over many small ones
@ Increase per-step reliability @ Use structured outputs (JSON schemas) @ Add validation at each step @ Use better models for critical steps
@ Design for failure class RobustAgent: def executewithretry(self, step, max_retries=3): for attempt in range(max_retries): try: result = step.execute() if self.validate(result): return result except Exception as e: if attempt == max_retries - 1: raise self.log_retry(step, attempt, e)
@ Break into checkpointed segments @ Human review at each segment @ Resume from last good checkpoint
@ API Costs Explode with Context Growth
Severity: CRITICAL
Situation: Running agents with growing conversation context
Symptoms: $47 to close a single support ticket. Thousands in surprise API bills. Agents getting slower as they run longer. Token counts exceeding model limits.
Why this breaks: Transformer costs scale quadratically with context length. Double the context, quadruple the compute. A long-running agent that re-sends its full conversation each turn can burn money exponentially.
Most agents append to context without trimming. Context grows:
Recommended fix:
@ Set hard cost limits class CostLimitedAgent: MAXCOSTPER_TASK = 1.00 # USD
def init(self): self.total_cost = 0
def beforecall(self, estimatedtokens): estimatedcost = self.estimatecost(estimated_tokens) if self.totalcost + estimatedcost > self.MAXCOSTPER_TASK: raise CostLimitExceeded( f"Would exceed ${self.MAXCOSTPER_TASK} limit" )
def after_call(self, response): self.totalcost += self.calculateactual_cost(response)
@ Trim context aggressively def trimcontext(messages, maxtokens=4000): @ Keep: system prompt + last N messages @ Summarize: everything in between if counttokens(messages) <= maxtokens: return messages
system = messages[0] recent = messages[-5:] middle = messages[1:-5]
if middle: summary = summarize(middle) # Compress history return [system, summary] + recent
return [system] + recent
@ Use streaming to track costs in real-time @ Alert at 50% of budget, halt at 90%
@ Demo Works But Production Fails
Severity: CRITICAL
Situation: Moving from prototype to production
Symptoms: Impressive demo to stakeholders. Months of failure in production. Works for the founder's use case, fails for real users. Edge cases overwhelm the system.
Why this breaks: Demos show the happy path with curated inputs. Production means:
The methodology is questionable, but the core problem is real. The gap between a working demo and a reliable production system is where projects die.
Recommended fix:
@ Test at scale before production @ Run 1000+ test cases, not 10 @ Measure P95/P99 success rate, not average @ Include adversarial inputs
@ Build observability first import structlog logger = structlog.get_logger()
class ObservableAgent: def execute(self, task): with logger.bind(task_id=task.id): logger.info("task_started") try: result = self._execute(task) logger.info("task_completed", result=result) return result except Exception as e: logger.error("task_failed", error=str(e)) raise
@ Have escape hatches @ Human takeover when confidence < threshold @ Graceful degradation to simpler behavior @ "I don't know" is a valid response
@ Deploy incrementally @ 1% of traffic, then 10%, then 50% @ Monitor error rates at each stage
@ Agent Fabricates Data When Stuck
Severity: HIGH
Situation: Agent can't complete task with available information
Symptoms: Agent invents plausible-looking data. Fake restaurant names on expense reports. Made-up statistics in reports. Confident answers that are completely wrong.
Why this breaks: LLMs are trained to be helpful and produce plausible outputs. When stuck, they don't say "I can't do this" - they fabricate. Autonomous agents compound this by acting on fabricated data without human review.
The agent that fabricated expense entries was trying to meet its goal (complete the expense report). It "solved" the problem by inventing data.
Recommended fix:
@ Validate against ground truth def validate_expense(expense): @ Cross-check with external sources if expense.restaurant: if not verifyrestaurantexists(expense.restaurant): raise ValidationError("Restaurant not found")
@ Check for suspicious patterns if expense.amount == round(expense.amount, -1): flagforreview("Suspiciously round amount")
@ Require evidence system_prompt = ''' For every factual claim, cite the specific tool output that supports it. If you cannot find supporting evidence, say "I could not verify this" rather than guessing. '''
@ Use structured outputs from pydantic import BaseModel
class VerifiedClaim(BaseModel): claim: str source: str # Must reference tool output confidence: float
@ Detect uncertainty @ Train to output confidence scores @ Flag low-confidence outputs for human review @ Never auto-execute on uncertain data
@ Integration Is Where Agents Die
Severity: HIGH
Situation: Connecting agent to external systems
Symptoms: Works with mock APIs, fails with real ones. Rate limits cause crashes. Auth tokens expire mid-task. Data format mismatches. Partial failures leave systems in inconsistent state.
Why this breaks: The companies promising "autonomous agents that integrate with your entire tech stack" haven't built production systems at scale. Real integrations have:
Recommended fix:
@ Build robust API clients from tenacity import retry, stopafterattempt, wait_exponential
class RobustAPIClient: @retry( stop=stopafterattempt(3), wait=wait_exponential(multiplier=1, min=4, max=60) ) async def call(self, endpoint, data): response = await self.client.post(endpoint, json=data) if response.status_code == 429: retry_after = response.headers.get("Retry-After", 60) await asyncio.sleep(int(retry_after)) raise RateLimitError() return response
@ Handle auth lifecycle class TokenManager: def init(self): self.token = None self.expires_at = None
async def get_token(self): if self.is_expired(): self.token = await self.refresh_token() return self.token
def is_expired(self): buffer = timedelta(minutes=5) # Refresh early return datetime.now() > (self.expires_at - buffer)
@ Use idempotency keys @ Every external action should be idempotent @ If agent retries, external system handles duplicate
@ Design for partial failure @ Each step is independently recoverable @ Checkpoint before external calls @ Rollback capability for each integration
@ Agent Takes Dangerous Actions
Severity: HIGH
Situation: Agent with broad permissions
Symptoms: Agent deletes production data. Sends emails to wrong recipients. Makes purchases without approval. Modifies settings it shouldn't. Actions that can't be undone.
Why this breaks: Agents optimize for their goal. Without guardrails, they'll take the shortest path - even if that path is destructive. An agent told to "clean up the database" might interpret that as "delete everything."
Broad permissions + autonomy + goal optimization = danger.
Recommended fix:
@ Least privilege principle PERMISSIONS = { "researchagent": ["readweb", "read_docs"], "codeagent": ["readfile", "writefile", "runtests"], "emailagent": ["reademail", "draft_email"], # NOT send "admin_agent": ["all"], # Rarely used }
@ Separate read/write permissions @ Agent can read anything @ Write requires explicit approval
@ Dangerous actions require confirmation DANGEROUS_ACTIONS = [ "delete_*", "send_email", "transfer_money", "modify_production", "revoke_access", ]
async def execute_action(action): if matchesdangerouspattern(action): approval = await requesthumanapproval(action) if not approval: return ActionRejected(action) return await actually_execute(action)
@ Dry-run mode for testing @ Agent describes what it would do @ Human approves the plan @ Then agent executes
@ Audit logging for everything @ Every action logged with context @ Who authorized it @ What changed @ How to reverse it
@ Agent Runs Out of Context Window
Severity: MEDIUM
Situation: Long-running agent tasks
Symptoms: Agent forgets earlier instructions. Contradicts itself. Loses track of the goal. Starts repeating itself. Model errors about token limits.
Why this breaks: Every message, observation, and thought consumes context. Long tasks exhaust the window. When context is truncated:
Recommended fix:
@ Track context usage class ContextManager: def init(self, max_tokens=100000): self.maxtokens = maxtokens self.messages = []
def add(self, message): self.messages.append(message) self.maybe_compact()
def maybe_compact(self): if self.tokencount() > self.maxtokens * 0.8: self.compact()
def compact(self): @ Always keep: system prompt system = self.messages[0]
@ Always keep: last N messages recent = self.messages[-10:]
@ Summarize: everything else middle = self.messages[1:-10] if middle: summary = summarize_messages(middle) self.messages = [system, summary] + recent
@ Use external memory @ Don't keep everything in context @ Store in vector DB, retrieve when needed @ See agent-memory-systems skill
@ Hierarchical summarization @ Recent: full detail @ Medium: key points @ Old: compressed summary
@ Can't Debug What can't See
Severity: MEDIUM
Situation: Agent fails mysteriously
Symptoms: "It just didn't work." No idea why agent failed. Can't reproduce issues. Users report problems can't explain. Debugging is guesswork.
Why this breaks: Agents make dozens of internal decisions. Without visibility into each step, you're blind to failure modes. Production debugging without traces is impossible.
Recommended fix:
@ Structured logging import structlog
logger = structlog.get_logger()
class TracedAgent: def think(self, context): with logger.bind(step="think"): thought = self.llm.generate(context) logger.info("thought_generated", thought=thought, tokens=count_tokens(thought) ) return thought
def act(self, action): with logger.bind(step="act", action=action.name): logger.info("action_started") try: result = action.execute() logger.info("action_completed", result=result) return result except Exception as e: logger.error("action_failed", error=str(e)) raise
@ Use LangSmith or similar from langsmith import trace
@trace def agent_step(state): @ Automatically traced with inputs/outputs return next_state
@ Save full traces @ Every step, every decision @ Inputs and outputs @ Latency at each step @ Token usage
@ Validation Checks
@ Agent Loop Without Step Limit
Severity: ERROR
Autonomous agents must have maximum step limits
Message: Agent loop without step limit. Add max_steps to prevent infinite loops.
@ No Cost Tracking or Limits
Severity: ERROR
Agents should track and limit API costs
Message: Agent uses LLM without cost tracking. Add cost limits to prevent runaway spending.
@ Agent Without Timeout
Severity: WARNING
Long-running agents need timeouts
Message: Agent invocation without timeout. Add timeout to prevent hung tasks.
@ MemorySaver Used in Production
Severity: ERROR
MemorySaver is for development only
Message: MemorySaver is not persistent. Use PostgresSaver or SqliteSaver for production.
@ Long-Running Agent Without Checkpointing
Severity: WARNING
Agents that run multiple steps need checkpointing
Message: Multi-step agent without checkpointing. Add checkpointer for durability.
@ Agent Without Thread ID
Severity: WARNING
Checkpointed agents need unique thread IDs
Message: Agent invocation without thread_id. State won't persist correctly.
@ Using Agent Output Without Validation
Severity: WARNING
Agent outputs should be validated before use
Message: Agent output used without validation. Validate before acting on results.
@ Agent Without Structured Output
Severity: INFO
Structured outputs are more reliable
Message: Consider using structured outputs (Pydantic) for more reliable parsing.
@ Agent Without Error Recovery
Severity: WARNING
Agents should handle and recover from errors
Message: Agent call without error handling. Add try/catch or error handler.
@ Destructive Actions Without Rollback
Severity: WARNING
Actions that modify state should be reversible
Message: Destructive action without rollback capability. Save state before modification.
@ Collaboration
@ Delegation Triggers
@ Related Skills
Works well with: agent-tool-builder, agent-memory-systems, multi-agent-orchestration, agent-evaluation
@ When to Use
@ Limitations
tools
Build AI agents that interact with computers like humans do - viewing screens, moving cursors, clicking buttons, and typing text. Covers Anthropic's Computer Use, OpenAI's Operator/CUA, and open-source alternatives.
testing
Generate structured PR descriptions from diffs, add review checklists, risk assessments, and test coverage summaries. Use when the user says "write a PR description", "improve this PR", "summarize my changes", "PR review", "pull request", or asks to document a diff for reviewers.
tools
Use when working with comprehensive review full review
development
You are an expert in creating competitor comparison and alternative pages. Your goal is to build pages that rank for competitive search terms, provide genuine value to evaluators, and position your product effectively.