How Can Agents Generate Enterprise Data Without Database Schemas?
Introduction
Training and evaluating tool-calling agents requires more than isolated question-and-answer examples. Agents must encounter realistic entities, state transitions, permissions, and multi-step workflows. In enterprise settings, however, the systems and records needed to create such data are often protected by privacy, legal, and commercial restrictions. Database schemas may be unavailable as well.
Synthetic data appears to be a natural solution, but conventional approaches face a difficult trade-off. Table-oriented synthesizers can model statistical patterns, yet they may produce records that violate relationships or business rules. Procedure-based generators can create valid records by following hand-written workflows, but they usually require extensive domain-specific authoring and may not reproduce the distribution of real business data.
The paper “Synthesis Through Simulation” proposes a different abstraction. Instead of asking a model to fill tables directly, it places an LLM agent inside a simulated enterprise environment and lets the agent generate data by calling APIs.
The central idea
In STS, the simulated environment exposes operations for creating, retrieving, or updating business objects. These APIs enforce policies and state-dependent rules. A generated object becomes part of the environment only if the corresponding operation is accepted. Validity is therefore determined by the same operational layer that defines what the simulated system allows.
This separates two responsibilities:
- The environment enforces validity. Relationships, permissions, state transitions, and other constraints are checked during interaction rather than reconstructed after free-form generation.
- The agent models distribution. The agent decides which operations to call, in what order, and how to combine them to create a varied population of entities and workflows.
- The method avoids schema access. STS’s Generalist Populator is designed to work without reading the underlying database schema. It learns what is possible through API interaction and feedback.
- The interface can be reused across domains. Domain-specific rules live in the simulated environments, while the population agent remains general-purpose instead of being rewritten for every business area.
What the reported evaluation shows
The authors evaluate the Generalist Populator across ten simulated environments. They report an average marginal fidelity of 0.88 and 100% constraint satisfaction across all environments, while the agent has no access to database schemas. Marginal fidelity measures how closely generated distributions match the target along individual dimensions; constraint satisfaction measures whether the resulting data complies with the environment’s rules.
The comparisons are also informative. Statistical synthesizers could not be applied to seven of the ten environments because they required seed data. A schema-privileged agent, despite receiving database-structure information, failed 82% of trajectories in the airline environment. That environment contains tightly coupled workflows, illustrating that knowing fields and tables does not necessarily provide an operational understanding of the system. Correct sequencing and state-dependent behavior can matter more than structural metadata.
Why it matters
STS reframes enterprise data synthesis as system simulation. This is particularly relevant for agent research because the output is not only a collection of records; it can also include interaction trajectories that resemble tool use in a business application. Researchers may be able to train and evaluate agents without exposing real enterprise data or internal database structures.
The approach does not eliminate the need for careful modeling. Its guarantees apply to the simulated environment, not automatically to the real organization. If the environment omits important rules or represents an unrealistic distribution of operations, the generated data can be internally valid while still being a poor proxy for production behavior. Exploration cost, API feedback design, and evaluation of joint rather than marginal distributions remain important questions.
The authors release the framework, all ten environments, and generated datasets. STS therefore offers both a practical testbed and a broader research direction: building controllable simulations in which agents learn enterprise behavior through interaction rather than through unrestricted table generation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...