An arXiv preprint reports that a retrieval method for AI agents posted the highest task score on both simulated benchmarks for each of six language-model backbones. It also used fewer mean environment steps than the Graph-of-Skills (GoS) baseline in all 12 model and benchmark settings.
Across the six backbones, the ScienceWorld score was 72.62 with GoS versus 80.50 with CaSKG. On ALFWorld, success was 80.01% with GoS versus 86.79% with CaSKG. These are aggregate descriptive comparisons, and no confidence intervals or significance tests were reported.
The method is built around link quality
CaSKG first builds a high-recall directed candidate graph from semantic, lexical, input/output and structural evidence. Here, high-recall means the system starts with many plausible directional links. Repair evidence is also used, and an optional large language model can refine candidate scores.
Selected directed edges are then examined with textual counterfactual probes. The probes remove the source skill, substitute a dissimilar skill and reverse the proposed order. These checks ask whether the proposed relationship remains supported when the procedure is changed in those ways.
The publication step keeps confirmed relations at full support, downweights uncertain relations, removes rejected relations and retains selected unvalidated candidates as weaker scaffold links. In effect, the graph can carry provisional structure without treating every proposed link as equally strong.
The comparison used shared ground rules
The evaluation used 140 ALFWorld ID-140 household-task episodes and 211 ScienceWorld U211 science-task episodes. The main comparison covered six backbones: MiniMax-M2.7, GLM-5.2, Kimi-K2.6, Qwen3.5-397B-A17B, DeepSeek-V4-Flash and GPT-5.6-Luna.
All methods used the same task cohort, prompts, evaluators, episode limits and environment interaction loop. CaSKG published its graph before evaluation, did not update it during an episode and did not add a separate online planner.
The advantage was broad, but not universal
On the reported environment-interaction measure, CaSKG's six-model mean was 15.29 steps versus 16.39 for GoS on ScienceWorld. On ALFWorld, it was 14.05 versus 15.96. The figures are averages across the six backbones, so they summarize the model set rather than any one system.
The pattern was broad across the ScienceWorld task mix, but it was not universal. Across all 24 U211 task types, CaSKG scored higher on 21, tied on one and scored lower than GoS on two. The representative trajectory examples in the paper were presented as illustrations of failure modes, not independent proof of the aggregate results.
A closer look at scale and selectivity
The result also held as the skill library was varied in a two-backbone comparison. At every tested library scale, CaSKG had higher reported success than GoS for MiniMax-M2.7 and Qwen3.5-397B-A17B. Because this analysis covered only two backbones, it is a focused check rather than a full test across the model set.
In a component comparison, Full CaSKG reported 73.57% success and 18.44 mean steps while publishing 3,292 of 9,937 candidate relations. Publishing all candidates reported 71.43% success and 18.74 mean steps, while publishing all 9,937 relations. In the full version, selective publication kept only part of the candidate set, while selected unvalidated links could remain as weaker scaffolds.
The evidence stays inside the simulation
The study used simulated ALFWorld and ScienceWorld environments, did not collect personal data or involve human subjects, and did not infer sensitive attributes.
The manuscript is identified as arXiv:2608.25500v1 [cs.AI], dated 26 Aug 2026. No funding source is reported in the supplied text or metadata.
Paper data and sources
Original title: CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
Authors: Zhiyuan Li, Linyuan Gao, Xuechun Ding et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text