Gan Jiang: A Scientific Agent That Learns from XRD Failures
Introduction
Powder X-ray diffraction (XRD) is widely used to identify crystalline phases, quantify mixtures, and follow structural changes in materials. Yet reliable analysis is rarely a single database lookup. It may require candidate matching, multiphase decomposition, parameter refinement, and physical validation. Overlapping reflections and differences between analysis packages make this workflow difficult to automate robustly.
The paper introduces Gan Jiang, a scientific agent designed to turn this workflow into an adaptive process. Rather than retraining a language model to store more domain knowledge, the system represents analytical experience as executable skills and revises those skills when an analysis fails.
Key points
- An integrated diffraction ecosystem. Gan Jiang works with XMatcher, XQueryer, XDecomposer, and WPEM. Together, these components support phase matching, information retrieval, multiphase decomposition, and physics-constrained whole-pattern modelling.
- Skill-level learning. After diagnosing a failure, the agent can revise both skill instructions and code. A revised skill is validated before it is reused, while the underlying language model and physical models remain unchanged.
- Evaluation with held-out data. Skills are selected using development data and frozen for held-out evaluation. The paper compares refinement performance across FullProf, GSAS-II, and PyWPEM.
- Scientific case studies. The reported applications include strongly overlapping reflections, quantitative analysis of a five-phase ancient Egyptian cosmetic, lattice evolution in an operating battery, and comparisons of atomic configurations in a disordered oxide catalyst.
- Results on DeltaXRDbench. Gan Jiang leads the evaluated methods for single- and multiphase identification on simulated and experimental data. Without supplied composition, the reported single-phase top-1 accuracies on MP500, RRUFF, and opXRD are 96.30%, 81.78%, and 40.83%, respectively.
Why it matters
The main contribution is not simply giving an agent access to more scientific software. It is the proposed loop of analysis, failure diagnosis, skill revision, validation, and reuse. In scientific computing, a convincing explanation is not enough: the conclusion should remain tied to measured diffraction data, physical constraints, and reproducible calculations.
The work also suggests that scientific agents may gain an important part of their capability from tool ecosystems and experience management, rather than from model scale alone. Several practical questions remain important, including how stable revised skills are on new samples, how often skills must be superseded, and whether selecting a stored skill is cheaper and more reliable than re-deriving an analysis. XRD provides a useful testbed because a bad procedure can ultimately be exposed through an incorrect physical phase assignment.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...