about
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents (arxiv.org)
57 points by omarsar 24 days ago | hide | past | pdf | 15 comments on HN

In plain words: Agents get a graph of linked steps that says what to do next, with a helper turning nearby steps into a hint at each decision. Starting from a bare skeleton, it edits itself using failed and successful runs, matching hand-built graphs and beating memory-based agents.

Abstract

Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order, and under which conditions. As trajectories lengthen, agents can lose track of their objectives, invoke tools out of order, and repeat unproductive actions. We introduce the Procedural Graph: just as a knowledge graph organizes factual knowledge into (entity, relation, entity) triplets for what-is questions, a Procedural Graph organizes procedural knowledge into (procedure, relation, procedure) triplets for what-to-do questions. At each decision step, the framework localizes the agent's active node, and a guidance model translates the surrounding subgraph into step-level situational guidance that biases the solver's next action without dictating it. The graph is self-evolving: an LLM refiner contrasts failed trajectories with successful ones and edits the graph's topology and attributes, committing edits that preserve or improve held-out validation performance while retaining rejected ones to discourage repetition. Starting from a minimal skeleton, the loop builds graphs that match or surpass hand-designed ones. It can also repair a flawed expert prior. Across multiple datasets, task types, and LLMs, the Procedural Graph delivers consistent gains over memory-based baselines, and self-evolution further improves performance without manual engineering.

Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan Ö. Arık
arXiv:2609.09153 · cs.AI, cs.CL, cs.MA · submitted Sep 8, 2026
abstract · pdf · html · 36 pages including references and appendices, 6 figures, 11 tables

add comment on HN
Also discussed: Sep 2026 (4 points, 0 comments)

I wonder if the graph part is a distraction (data representation / syntax), and once you step back, if/how this relates to the wider concept of 'durable task lists' used in this kind of structured dynamic planning.

Ex: Durable task lists generally use more textual representations, eg, a hierarchical text list that supports named references -- so a graph. Likewise, they're mutable, contain statuses, etc. They're fairly popular and AI models at this point have internalized them at this point.

I mean, sentences themselves are graphs of data and context with connectors of various types. Embeddings within these models help form the links to the underlying semantic concepts.

Directly representing things as a graph would likely reduce some of the 'translation overhead', done correctly.

Very much agreed in the former

I'm not so sure on the latter -- that introduces a bunch of tool calls, while the text patches are much closer to the semantic space imo and easy for harnesses

If we rephrase this as instruction following alignment, what is the concern here in practice vs a skill? (Which a model eventually internalizes)

It doesn't detract from the research - as a paper, it shows more crisply the structure is useful. I'm just not sure how necessary the encoding, and given durable tasks, novel the insight. Is there new alpha here somewhere, esp given the similarity?

Maybe I missed in when scanning the paper, but did they compare to a simple self modifying harness, e.g. instructing Pi to update some code or skill docs based on the results?
Okay how are Add nodes created? Does an LLM come up with the name, guidance and edges for each node or is it a bespoke transformer model?
No idea what this site is. Paper is here: https://arxiv.org/abs/2609.09153
Thanks, we've updated the link.
The distinction between a procedural graph and a static workflow seems important: self-editing topology can capture reusable strategy, but it also makes regressions harder to localize. I’d be curious whether the refinement loop treats a successful trajectory as sufficient evidence, or uses counterexamples and held-out tasks to avoid encoding a brittle shortcut. A practical evaluation might report graph churn and rollback frequency alongside task success, since a graph that keeps growing could be trading inference cost and auditability for a small gain. The explicit entity–relation–procedure representation also seems like a promising place to attach permissions or provenance to tool calls.
What if open-ended agents are overkill for 99% of problems? Let's just take that premise for a second. Most organizations want to follow "best practices" and train their employees to do so. Hiring a genius and giving him total freedom to complete every task is not what companies usually want for MOST things. They want repeatability, reliability, predictability. Especially if they are regulated.

I'm going to drop a bomb over here: what if agents are the root of all our problems in AI safety, cost, and even adoption by regulated organizations? I really do believe this. Here is what I argue:

https://safebots.ai/agents.html

https://safebots.ai/kimi.html

And here is my overall thesis:

https://safebots.ai/thesis.html

If you do manage to read (or skim) that, I welcome any questions, comments or rebuttals.

Don't post generated text or AI-edited text. HN is for conversation between humans.
what if it's not generated text, but autism or god forbid, a german.
Achtung!
your load-bearing thesis is probably interesting but it seems i can't read AI written text anymore -- or maybe i just need some more coffee.
im happy for you tho, or sorry that happened
I also believe that, but also can’t follow LLM copy.