Recover from tool failure
Bonus: recover from a tool failure
Section titled “Bonus: recover from a tool failure”Time: 3–5 minutes · No edits required
The weather tool throws an actionable error when a date is outside Open-Meteo’s forecast window. Trigger that path deliberately:
npm run deploysleep 5TIMEOUT_S=200 npm run smoke -- https://field-trip-agent.<subdomain>.workers.dev failure-live-1 \ "We're planning a Lisbon offsite on 2099-03-01 for 8 people. Call the weather tools and explain the result."TIMEOUT_S=200 npm run smoke -- https://field-trip-agent.<subdomain>.workers.dev failure-live-1 \ "Change the offsite dates to <START> through <END>. Save the corrected brief and try the weather tools again."Save your workers.dev subdomain to open this test ↗https://field-trip-agent.<subdomain>.workers.dev/?id=failure-live-1
This opens the same conversation as the smoke test. If you repeat this checkpoint, before copying its commands.
Start npm run dev in Terminal A and run in Terminal B:
TIMEOUT_S=200 npm run smoke -- http://localhost:5173 failure-1 \ "We're planning a Lisbon offsite on 2099-03-01 for 8 people. Call the weather tools and explain the result."TIMEOUT_S=200 npm run smoke -- http://localhost:5173 failure-1 \ "Change the offsite dates to <START> through <END>. Save the corrected brief and try the weather tools again."Open local chat in your browser ↗http://localhost:5173/?id=failure-1
This opens the same conversation as the smoke test. If you repeat this checkpoint, before copying its commands.
Verification gate
Prove it works
save_trip_briefandgeocode_citysucceed.get_forecastdisplays an error rather than invented weather.- The agent catches the tool outcome and explains the 16-day limit.
- The second request updates the saved dates and gets a real forecast in the same conversation.
- The first
execute_tool get_forecastspan records an error while the enclosing agent turn can still complete. - The corrected forecast succeeds in the next turn’s trace.
Compare both turns in the Agents dashboard. The trace waterfall separates a failed tool operation from the final status of the complete agent response.
No code reset is needed. The saved brief now contains the corrected dates. Use a fresh conversation ID for your next experiment.