Servicenow Coreai Builds Autosynthdata for Enterprise Agent Training

ServiceNow CoreAI introduced AutoSynthData, a pipeline that turns a target model's observed failures into validated synthetic training tasks for enterprise agents, using a stronger teacher model to demonstrate correct behavior. In the EnterpriseOps Gym Hybrid environment, 2,000 generated samples improved mean Pass@1 by 7.2 percentage points. A second ITSM run lifted mean Pass@1 from 18.77% to 27.18%.

Published: October 2, 2026 By David Kim, AI & Quantum Computing Editor AI Author Category: Agentic AI

David focuses on AI, quantum computing, automation, robotics, and AI applications in media. Expert in next-generation computing technologies.

Servicenow Coreai Builds Autosynthdata for Enterprise Agent Training

Executive Summary

  • ServiceNow CoreAI introduced AutoSynthData, a pipeline that converts a target model's observed failures into new synthetic training tasks for enterprise agents, using a stronger teacher model to demonstrate successful behavior (Hugging Face).
  • The system generates tasks as a combination of system specification, user prompt and verifier, and requires every candidate to pass execution, positive and negative verification, solver-based difficulty measurement and bounded repair before it enters a training set (Hugging Face).
  • In the EnterpriseOps Gym Hybrid environment, 2,000 synthetic samples generated in about 18 hours improved mean Pass@1 by 7.2 percentage points, a 35% relative gain, and raised verifier success from 63.01% to 68.55% (Hugging Face).
  • In the ITSM environment, 1,994 samples generated in 66 hours lifted mean Pass@1 from 18.77% to 27.18%, extending the reported result to a second enterprise domain (Hugging Face).

Key Takeaways

  • AutoSynthData treats failure patterns, not generic instruction volume, as the input that determines what gets generated next.
  • Quality control operates at two levels: per-sample gates covering feasibility, soundness and difficulty, and batch-level review of coverage, diversity and redundancy.
  • Reported gains come from supervised fine-tuning experiments in a controlled benchmark; the article presents reinforcement learning as an untested extension.
  • The pipeline's throughput varied widely between the two reported runs, and the authors attribute the slower ITSM run partly to a larger teacher model and the absence of later optimizations.

ServiceNow CoreAI's AutoSynthData Pipeline Explained

The problem the ServiceNow CoreAI team describes is specific. Enterprises need agents that operate inside their own systems, rules and data states, and a broadly capable model can still fail on a particular workflow, tool combination or policy constraint. A single observed failure carries information, but post-training requires many new tasks that exercise the same capability in different situations. Those tasks must be executable in the environment, resemble work a user would plausibly request, and be checkable.

AutoSynthData structures generation around three components: a system specification that defines constraints and any task-specific initialization, an agent-facing user prompt, and a verifier that judges whether the resulting trajectory completed the task. The article sets explicit properties for each. Prompts must be feasible, realistic and difficult enough to expose a current weakness. Verifiers must be consistent with the prompt and environment state, sound enough to reject invalid trajectories, and complete enough to accept valid solutions rather than one reference path.

The loop begins with diagnostic evaluation runs of both the target model and a stronger teacher. The team distills findings into sanitized capability specification cards that describe the capability under test, the tools and workflow structure involved, where the target fails and the teacher succeeds, the properties a correct final state must satisfy, and the dimensions that can vary without changing the capability. The generator receives these cards, not the original prompts, entities, trajectories or verifier details.

Task Validation and Curriculum Design in AutoSynthData

Generation runs in two phases. The target phase produces a core set of vetted samples built around identified weaknesses, with workers generating tasks in parallel and each candidate passing validation, execution, solver evaluation and repair. The multiply phase expands accepted target samples into novel variants with their own requests, states, entity configurations, reference trajectories and verifiers. A multiplied sample cannot seed another multiplied sample, a constraint the authors describe as anchoring expansion to the vetted target set and limiting drift.

Difficulty is measured rather than assumed. In the configuration used here, the pipeline favors tasks the target model solves on no more than one of three trials while the stronger solver succeeds on at least two of three. Candidates then face positive verification, which executes the reference trajectory and checks the resulting state against the verifier, and negative verification, which mutates parts of the expected outcome to confirm those states no longer pass. Failed candidates go to a critic that looks for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic or a mismatch with the intended capability, after which targeted repairs are attempted within a fixed retry limit.

Related: Waymo Faces Transparency Scrutiny Amid Remote Worker Concerns in 2026

Batch-level review addresses a different failure mode. A generation batch may overrepresent a few easy task families, miss a capability, or consume budget on a low-yield pattern. A meta-review examines accepted and rejected samples and generation behavior, asks which families are overrepresented and which capability dimensions are missing, and guides changes to generation strategy. The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions and directs work toward gaps.

EnterpriseOps Gym Results for AutoSynthData

The reported evidence comes from EnterpriseOps Gym, cited as Malay et al., 2026. In the Hybrid domain, Gemma-4-26B-A4B-it served as the target model and Qwen3.8-27B as the teacher. AutoSynthData generated 2,000 synthetic training samples in about 18 hours. After fine-tuning and checkpoint evaluation, the best checkpoint was epoch 5. That checkpoint improved mean Pass@1 by 7.2 percentage points, a 35% relative improvement, raised verifier success from 63.01% to 68.55%, and closed 59% of the original Pass@1 gap between the target and the reference model. The authors state the training tasks were newly generated from capability specifications and that the generator did not receive the original evaluation tasks.

For deeper context, see our Climate Tech analysis: "Nyobolt Series C 2026: Cambridge Battery Firm Hits $1B at 60M Round".

In the ITSM domain, the same target model was paired with DeepSeek-V4.1-Flash as teacher. The run produced 1,994 samples over 66 hours. The article attributes the longer generation time primarily to the larger teacher model and to preceding pipeline optimizations that improved throughput in the later Hybrid run. Synthetic SFT raised mean Pass@1 from 18.77% to 27.18%.

These are benchmark results in a controlled setting, reported by the team that built the pipeline. They show improvement in the specific environments tested, not a general claim about enterprise agent performance. The article notes that experiments focused on supervised fine-tuning and that the same mechanism could support reinforcement learning, which the team plans to test. That remains a stated intention rather than a demonstrated result.

Additional coverage: NVIDIA CEO Jensen Huang CMU Speech 2026: AI Industrial Era Message

AutoSynthData Signals

Entity Recent Focus Geography Source
ServiceNow CoreAI AutoSynthData pipeline that turns target model failures into validated synthetic training tasks for enterprise agents Not specified in source Hugging Face
AutoSynthData Two-phase task generation with target and multiply phases, sample-level verification and batch-level meta-review Not specified in source Hugging Face
EnterpriseOps Gym Hybrid and ITSM environments used to generate tasks, fine-tune checkpoints and evaluate results Not specified in source Hugging Face
Gemma-4-26B-A4B-it Target model in both the Hybrid and ITSM experiments Not specified in source Hugging Face
Qwen3.8-27B Teacher model for the Hybrid experiment, providing demonstrations for generated tasks Not specified in source Hugging Face
DeepSeek-V4.1-Flash Teacher model for the ITSM experiment, cited as a reason generation took longer Not specified in source Hugging Face

What This Means for Practitioners

For teams building or buying enterprise agents, the operational message is that evaluation failures can be treated as a generation target rather than a scorecard item. Practitioners running agents against internal systems should ask whether their failure data is captured in enough detail to describe the capability involved, the tools and workflow structure, and the properties a correct final state must satisfy. The article's quality gates also imply that accepting synthetic tasks without executing them is risky: feasibility, verifier soundness and difficulty calibration all required machine checks here, and weak verifiers rewarded wrong states. Budget planning should account for generation time and teacher cost.

Hugging Face Implementation Risks

The pipeline depends on a stronger teacher model being available and correct. Where the teacher fails or its demonstrations encode a particular path, verifier completeness is tested, and the article explicitly warns that an overly restrictive verifier can penalize valid solutions while a lax one can reward incorrect behavior. Generation cost is material: the ITSM run took 66 hours for 1,994 samples, and the authors attribute part of that to a larger teacher and unoptimized throughput, indicating that cost and latency depend heavily on configuration choices.

Results are limited to two benchmark domains in one environment, produced and reported by the team that built the system. The article does not provide independent replication, does not report reinforcement learning outcomes, and does not claim generalization beyond EnterpriseOps Gym. The multiply phase's no-seeding rule constrains drift but also constrains how far the dataset can extend from vetted targets.

Editorial independence disclosure: this article was written independently from the underlying research, and the only external source used is the canonical article linked below. Source note: AutoSynthData: Generating Training Data for Enterprise Agents, published on Hugging Face on October 2, 2026.

About the Author

DK

David Kim AI Author

AI & Quantum Computing Editor

David focuses on AI, quantum computing, automation, robotics, and AI applications in media. Expert in next-generation computing technologies.

David Kim is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →

About Our Mission Editorial Guidelines Corrections Policy Contact

Frequently Asked Questions

What is AutoSynthData?

AutoSynthData is a pipeline built by ServiceNow CoreAI that converts a target model's failures into new synthetic training tasks for enterprise agents. It uses a stronger teacher model's successes to characterize solvable behavior, then generates and validates new tasks that exercise the same capabilities, shifting the curriculum as the model improves.

How does AutoSynthData validate generated tasks?

Every candidate must be executable in the target environment, pass positive verification that its intended solution works, pass negative verification that incorrect outcomes fail, and clear solver-based difficulty measurement. Failed candidates go to a critic for targeted repair within a fixed retry limit. Batch-level meta-review then checks coverage, diversity and redundancy across accepted samples.

What results did AutoSynthData report in EnterpriseOps Gym?

In the Hybrid domain, 2,000 synthetic samples generated in about 18 hours improved mean Pass@1 by 7.2 percentage points, a 35% relative gain, and raised verifier success from 63.01% to 68.55%. In the ITSM domain, 1,994 samples generated in 66 hours raised mean Pass@1 from 18.77% to 27.18%. These are benchmark results reported by the team that built the pipeline.

Which models were used in the experiments?

Gemma-4-26B-A4B-it served as the target model in both the Hybrid and ITSM experiments. Qwen3.8-27B was the teacher model for Hybrid, while DeepSeek-V4.1-Flash was the teacher for ITSM. The article attributes the longer ITSM generation time partly to the larger teacher model and to pipeline optimizations added before the later Hybrid run.

Does AutoSynthData support reinforcement learning?

The source states that experiments focused on supervised fine-tuning, and that the same mechanism could support reinforcement learning by generating tasks that challenge the current policy and moving the generation target as the policy updates. The team says it plans to test this beyond SFT. The article does not report reinforcement learning outcomes.