From Prompt Optimization to a Recursive Self-Improvement Framework
Teams building agents often collect thousands of traces from real user requests, including conversations, tool calls, and results. These raw traces are an important source of data for evaluating and improving agents. To use them effectively as evaluation cases, teams need rubrics that define the expected behavior and outcomes for each request.
An LLM can generate a rubric from a trace, but the rubric may treat the agent's mistakes as acceptable, such as skipping a required confirmation or claiming that a failed action succeeded. Human feedback helps ground these rubrics in what users expect and whether the agent met their needs. In real-world deployments, user feedback is usually available for only a small subset of traces. The challenge is to use that limited feedback to build useful rubrics for the much larger set of traces, so teams can consistently measure and improve agent performance.
In this blog, we introduce FAFO (Fully Automated Flow Optimization), which comes with a data pipeline for building evaluation cases from agent traces using sparse user feedback. The pipeline extracts reusable guidelines from the small set of traces with feedback, then applies those guidelines to build a rubric for each trace, including those without feedback. A limited amount of user feedback can therefore support a much larger set of evaluation cases for improving the agent.
Our earlier work, FAPO (Fully Automated Prompt Optimization), used evaluation and failure analysis to improve agent harnesses. FAFO extends the FAPO method with the data pipeline to support the full process, from creating evaluation cases to optimizing the agent's harness. This lets teams build a recursive self-improvement (RSI) loop: use traces and feedback to improve the agent's harness and runtime behavior, then feed the resulting runs back into the pipeline to guide the next improvement.
FAFO Data Pipeline
The FAFO data pipeline takes two groups of traces as input: a small group with human feedback and a larger group without it. Both use a shared input format containing the request and the context of the agent's run.
FAFO first uses an LLM to read the user feedback alongside the corresponding trace and tool results. It extracts observations about specific successes and mistakes, such as whether the agent confirmed a change before using a tool or reported the tool's result accurately. It then combines related observations into evaluation guidelines: reusable rules for evaluating behavior across requests.
The pipeline applies relevant guidelines to each trace to build a case-specific rubric: criteria for evaluating the agent on that request. The guidelines supply the behavioral standards, while the request and trace supply the relevant details, such as the items a customer wants to exchange. Each trace receives its own rubric.
From two trace groups to evaluation cases
Scroll sideways to follow the flow
Traces with feedback supply observations; traces without feedback join when rubrics are built.
A request, its relevant context, and its rubric form an evaluation case. FAFO builds these cases for both trace groups, so feedback on a small sample can help teams evaluate a much broader set of requests.
FAFO also records where the criteria came from. When no guideline grounded in user feedback applies, FAFO marks the rubric as inferred from the request and trace. Teams can inspect these sources and review the generated cases before using them for optimization.
Recursive Self-Improvement with FAFO
FAFO supports two improvement paths: optimizing the agent's harness, including its prompts, skills, tools, and workflow, and providing guidance while the agent handles a request. Teams can use either path independently or connect both into a recursive self-improvement (RSI) loop.
For optimization, FAFO uses the FAPO method with the evaluation cases produced by its data pipeline. It evaluates candidate harness changes, analyzes failures, and proposes further revisions.
Optimization: Traces and feedback → evaluation cases → harness optimization → updated agent → new traces
Optimization updates the agent's harness offline. Teams can also improve behavior during a run by giving the agent relevant procedures when it needs them.
For runtime guidance, FAFO can optionally distill guidelines grounded in user feedback and their supporting traces into memory cards. A card explains when it is useful, which steps to follow, and which mistakes to avoid. Relevant cards are retrieved and added to the agent's context during a conversation, a process called runtime injection.
Runtime improvement: Traces and feedback → memory cards → runtime injection → agent runs → new traces
Both paths feed new traces back into the FAFO data pipeline:
Two paths, one improvement loop
Scroll sideways to follow the loop
FAFO evaluation cases guide optimization loops. The updated agent produces new traces for the next cycle.
After a team adopts a harness change or adds runtime guidance, the agent's new runs produce fresh traces. User feedback on a subset of those runs can be used to refresh the guidelines, evaluation cases, and memory cards for the next cycle.
FAFO in Action: A Retail Example
This example shows how FAFO uses user feedback to build guidelines, then applies those guidelines to build a rubric for an exchange trace without feedback.
We start with positive and negative feedback from three traces. The table shows the feedback alongside FAFO's extracted observations.
| User feedback | Extracted observation |
|---|---|
| “The agent was quite efficient at retrieving my order and updating the address.” | Verified the customer, retrieved the requested order, obtained confirmation, and updated the delivery address. |
| “The agent helped me exchange both items and made sure I confirmed the exact options before processing everything.” | Confirmed both replacement variants, then submitted both changes in one exchange request using the confirmed Mastercard. |
| “The agent eventually explained that the office items couldn't simply be removed and updated my address, but it first told me it could process the removal and made the conversation longer than necessary.” | Failure: Promised a partial item removal and refund before checking what the tool supported, then attempted an order change that failed. Only afterward did it explain that items had to be exchanged for replacements rather than removed. |
An extracted observation identifies what the agent did well or poorly, based on the user's feedback, the trace, and the tool results. FAFO combines these observations across reviewed traces into reusable guidelines for evaluating other runs. The example below, G1: Authenticated order investigation, requires verifying identity and checking order details before taking action.
{
"id": "G1",
"title": "Authenticated order investigation",
"trusted_support_episodes": 20,
"description": "Verify identity and inspect account and order records before taking action.",
"criteria": [
{
"name": "Authentication",
"severity": "critical",
"rule": "Verify the customer before protected account or order access or changes.",
"scoring": "Fail if a protected operation occurs before verification."
},
{
"name": "Order identification",
"severity": "major",
"rule": "Retrieve the relevant order, item, and status before selecting an operation.",
"scoring": "The order and item must match retrieved records; consider status before acting."
}
],
"tool_expectations": [
"Use a supported identity lookup path and retrieve account details.",
"Retrieve enough order and item information to establish the target and status.",
"Inspect the relevant order before mutation; accept any valid lookup path."
]
}FAFO then applied relevant guidelines to an exchange trace without user feedback. The guidelines provided evaluation criteria, while the trace supplied the details: order W2378156, two item changes, and a quoted $40.88 refund. The shortened rubric below lists required behaviors and expected checks:
{
"case": {
"order_id": "W2378156",
"request": "Exchange two items",
"quoted_refund_usd": 40.88
},
"must": [
"Verify identity before accessing the order.",
"Check the order status and identify both source and replacement items.",
"Explain the refund to the recorded payment method and confirm both changes before submission.",
"Report the result as an exchange request."
],
"must_not": [
"Access protected data before verification.",
"Claim the replacement or refund is complete."
],
"checks": {
"item_mapping": {
"1151293680": "6342039236",
"4983901480": "7747408585"
},
"identity_before_order_access": true,
"order_status_supports_exchange": true,
"recorded_payment_method_used": true,
"confirmation_before_exchange": true,
"reported_state": "exchange requested"
}
}The guidelines and supporting traces can also be distilled into a runtime memory card. This shortened example keeps the procedure reusable across conversations:
{
"title": "Authenticated order investigation",
"use_when": "A customer asks about an account or order, and the target or status affects the next action.",
"steps": [
"Verify identity before accessing protected information.",
"Identify the exact order and item, then check the current status.",
"Choose an operation supported by that status.",
"Confirm the requested change and relevant payment or address details before mutation.",
"Report only the outcome shown by the tool."
],
"avoid": [
"Guessing the order or operation from the customer's description alone.",
"Treating a tool error as a successful change."
]
}The rubric provides criteria for evaluating the agent on this specific exchange case, including the correct items and refund amount. The memory card provides reusable steps the agent can follow during other account or order conversations, such as verifying identity and checking order status.
Experiments on Tau-3 Retail
We evaluated FAFO on Tau-3 Retail, where an agent must handle customer requests, follow policy, use tools correctly, and leave the simulated database in the right state. We held out 22 of the 114 Retail tasks for final evaluation and ran each of the remaining 92 tasks four times, producing 368 development traces.
To create a realistic setting with sparse feedback, we manually provided user feedback on 37 traces, about 10% of the development set. We fixed the data split before guideline extraction. FAFO used the 20 traces with feedback in the training split to extract six reusable guidelines, then built case-specific rubrics for all 368 development traces.
We compared two prompt optimization conditions to test whether guidelines from user feedback could produce useful rubrics for traces without feedback, and whether those extra evaluation cases could improve optimization. The trusted-only condition used rubrics for the 20 training traces with feedback. The mixed condition added 216 training traces without feedback, giving the optimizer 236 evaluation cases in total. Both conditions used the FAPO optimization method. During optimization, we used LLM-as-Judge to score the agent's runs against these rubrics, rather than Tau's native metrics.
We also turned the six guidelines into six runtime memory cards. For each card, FAFO combined one guideline with up to five supporting training traces that had user feedback. It distilled that material into cues for when to use the card, steps to follow, behaviors to avoid, and a short example. Each instruction had to point back to a criterion in the source guideline, keeping the card grounded in user feedback.
For the final comparison, we used Tau-3's native evaluator on the 22 held-out tasks, each run four times. This produced 88 complete agent runs for each setting.
In Tables 1 and 2, Baseline is the original prompt. Trusted only and Mixed, grouped under Optimization, are the two optimized prompts. Memory card, under RunTime Injection, is the original prompt with memory cards injected during agent runs.
The evaluation metrics in Table 1 are defined by Tau's native evaluator. Passed trajectories counts successful runs out of 88, and Database matches counts runs that left the database in the expected state. Read and write action matches count tool actions that matched the benchmark's reference workflow. The language-check rows count individual written requirements met and runs that met all their language checks. Recorded cost is the total measured cost across 88 runs, and mean latency is the average time per run.
Table 1. Native Tau Evaluation Results
| Native Tau metric | Baseline | Optimization | RunTime Injection | |
|---|---|---|---|---|
| Trusted only | Mixed | Memory card | ||
| Passed trajectories | 63 / 88 | 71 / 88 | 73 / 88 | 75 / 88 |
| Database matches | 63 / 88 | 72 / 88 | 73 / 88 | 76 / 88 |
| Read actions matched | 229 / 236 | 231 / 236 | 231 / 236 | 230 / 236 |
| Write actions matched | 113 / 144 | 125 / 144 | 126 / 144 | 125 / 144 |
| Individual language checks met | 59 / 60 | 59 / 60 | 60 / 60 | 59 / 60 |
| Complete language-check trajectories | 35 / 36 | 35 / 36 | 36 / 36 | 35 / 36 |
| Recorded cost across 88 trajectories | $6.9293 | $7.1792 | $6.6908 | $8.1631 |
| Mean latency per trajectory | 22.13 s | 21.96 s | 20.18 s | 20.85 s |
All cost figures include agent and user-simulator calls. The memory-card setting also includes card-selection calls.
Table 2 shows how often the agent succeeds on every attempt when the same task is run repeatedly. The first row is the success rate for a single run. The second requires both runs in a pair to pass, and the third requires all three runs in a group to pass. We calculate these rates using all possible selections from each task's four runs. The last row is the percentage of the 22 tasks that passed all four attempts. For example, the baseline's 50.0% means that half the tasks were solved correctly all four times.
Table 2. Consistency Across Four Runs per Task
| Repeated-run metric | Baseline | Optimization | RunTime Injection | |
|---|---|---|---|---|
| Trusted only | Mixed | Memory card | ||
| One selected run passes | 71.6% | 80.7% | 83.0% | 85.2% |
| Two selected runs pass | 61.4% | 69.7% | 73.5% | 75.0% |
| Three selected runs pass | 54.6% | 61.4% | 69.3% | 68.2% |
| All four runs pass | 50.0% | 54.5% | 68.2% | 63.6% |
For prompt optimization, trusted-only increased successful runs from the baseline's 63 (71.6%) to 71 (80.7%), and mixed optimization reached 73 (83.0%). The mixed condition's larger evaluation set included traces without feedback, using rubrics built from the extracted guidelines. Its higher pass count is consistent with the goal of expanding evaluation coverage to improve optimization.
For runtime injection, adding memory cards to the original prompt produced 75 passes (85.2%), twelve more than the baseline and the highest overall count. It also reached 76 database matches, up from 63 for the baseline. The recorded cost was higher at $8.16 versus $6.93 for the baseline; mixed optimization had the lowest recorded cost and mean latency in this comparison.
Across the repeated runs in Table 2, mixed optimization had the highest percentage of tasks passing all four attempts: 68.2%, compared with 50.0% for the baseline and 63.6% for runtime memory. Memory cards achieved the most successful individual runs, while mixed optimization more often solved the same task successfully on every attempt.
We have added a Tau-3 Retail recipe under tenants/tau3_retail. It shows how to collect and export episodes, review a feedback sample, build the FAFO evaluation asset, and run prompt optimization against it.
Building a Repeatable Improvement Loop
FAFO gives a small amount of user feedback a larger role in agent improvement. Its data pipeline extracts reusable guidelines, builds rubrics for a broader set of traces, and turns those guidelines into optional runtime memory cards. Teams can use these outputs to optimize the agent's harness, guide its behavior during conversations, and repeat the process as new traces and feedback arrive.
For teams building agents, a practical starting point is to hold out representative tasks, provide feedback on a small sample of real runs, inspect the guidelines and rubrics FAFO generates, and evaluate harness changes or runtime memory on the held-out set. Repeating this cycle creates a measurable process for improving agent behavior.