Workshop date and room confirmed
FAGEN will take place on July 10 in Grand Ballroom 104-105 at COEX Convention & Exhibition Center, Seoul, South Korea.
View key datesReproducible Triggers, Trace Diagnostics, and Verified Fixes
FAGEN will take place on July 10 in Grand Ballroom 104-105 at COEX Convention & Exhibition Center, Seoul, South Korea.
View key datesFollow @FAGENWorkshop for deadline reminders, accepted-paper highlights, and program updates.
Follow @FAGENWorkshopThe submission portal is now live. Submission deadline May 11 (AOE); notifications by May 25.
Open OpenReviewReliability has been studied in ML for a long time, mostly through robustness benchmarks, adversarial evaluation, and red-teaming on chat-style language models. Foundation-model agents push the question somewhere harder. An agent run goes for hundreds of steps, each step depending on tool calls and memory writes from the steps before it. When the run breaks, it rarely breaks at the obvious moment. A bad assumption at step 3 quietly contaminates step 50, and by step 200 the agent has been wrong for a while without noticing. It might have spent its budget on the wrong subtask. It might be reading from memory it polluted itself. Or it landed on an answer at step 12 and spent the rest of the run defending it.
FAGEN is a place to take these failures seriously. The workshop is organized around four kinds of contributions. Definitions matter: what does "failure" actually mean here, beyond the loose way the term gets thrown around? Reproducible triggers matter at least as much. We want the smallest setup that breaks the agent the same way every time, so other groups can build on the case. Diagnostics should look at the trace itself, not just the final score, because final-score evaluation hides almost everything interesting. And the fixes worth presenting are the ones that admit what they cost in latency, in capability, or in how well they generalize.
Format
Submit on OpenReview
Topics of Interest
We welcome submissions on:
Well-documented negative results are in scope when the analysis is careful and the lesson transfers.
Operational definitions, triggering preconditions, minimal reproductions, composable failure primitives, and falsifiable mechanistic hypotheses.
Long-horizon evaluation protocols, interpretable process metrics, counterfactual tests, and logging tools that expose failures beyond terminal success.
Mitigations, recovery strategies, tool and memory interface improvements, reward and budget design, and repair mechanisms with verifiable trade-offs.
Workshop date
July 10, 2026
In person
GRAND BALLROOM 104-105
COEX Convention & Exhibition Center, Seoul, South Korea
Submission deadline
11
AOE · May 11, 2026 (AOE)
Notification date
25
AOE · May 25, 2026 (AOE)
Opening
Keynote Speech — Maarten Sap (online)
Keynote Speech — Greg Durrett (online)
Keynote Speech — Andrea Zanette
Oral Presentations
Coffee break + Poster Session 1
Keynote Speech — Nouha Dziri (online)
Keynote Speech — Yu Su (online)
Lunch
Keynote Speech — Bo Li
Keynote Speech — Rishi Bommasani
Oral Presentations and Special Invited Talk
Coffee break + Poster Session 2
Keynote Speech — Iryna Gurevych (online)
Keynote Speech — Samy Bengio
Sponsor Section
Best Paper + Closing
* In-person posters may be presented flexibly at either of the two poster sessions.
Reach out for submissions, sponsorship, speaker logistics, or collaboration.