Ai2 Replaces Priority GPU Scheduler With Time Budgets
Ai2's AI Infrastructure team replaced a priority-based GPU scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract. Over a 30-day test period, teams received 98% of the GPU hours they were owed while cluster occupancy held at 98%. Ai2 reports the change cut debug workload queue waits and reduced repairs needing a human in the loop by 74%.
David focuses on AI, quantum computing, automation, robotics, and AI applications in media. Expert in next-generation computing technologies.
Executive Summary
- Ai2's AI Infrastructure team replaced a priority-based GPU scheduler with a system combining GPU time budgets, hierarchical fair-share allocation and a time-slicing contract across clusters of NVIDIA H100, B200 and B300 GPUs ranging from 88 to 1024 units, Hugging Face.
- Over a 30-day test period following a cluster-by-cluster rollout that began at the end of July, teams received 98% of the GPU hours they were owed, with 13 of 15 team allocations at 95% or more and the worst case at 90%, Hugging Face.
- Cluster occupancy held steady at 98% before and after the change while demand exceeded capacity by 2-3x in both periods; 18% of delivered GPU time was unallocated, Hugging Face.
- Debug workload p90 queue wait time fell from 2 hours to 30 seconds in production, and repairs requiring a human in the loop dropped 74%, Hugging Face.
- The write-up documents costs as well as gains: interactive sessions were capped by the 8-hour protected runtime, prompting two new roadmap projects, Hugging Face.
Key Takeaways
- Ai2 shifted the debate over GPU time from case-by-case operational negotiation to an administrative budgeting process in which every protected request must be funded by an allocation or it is not shielded from preemption.
- The scheduler separates allocated occupancy, which is charged to a budget and protected for a minimum runtime, from unallocated occupancy, which is free but preemptible from the outset; that split is how the cluster stayed near full occupancy.
- Simulation preceded rollout. Ai2 built a scheduler simulator to test lookback-window length and minimum-runtime caps, and production outcomes overperformed its hand-built predictions.
- Not every workload improved. Long-held interactive sessions lost protected runtime, and Ai2 is investigating capacity fragmentation that could raise queue waits for the largest workloads.
Hugging Face Report Details Ai2's Budget-Based Scheduler Swap
The verified development is a scheduler replacement, not a hardware purchase. Ai2's AI Infrastructure team states it moved from a priority-based scheduler to a system built on three components: GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract that requires each workload to declare a minimum runtime in exchange for protected occupancy. The team describes the task as a pyramid of four metrics: availability, occupancy, impact, and utilization, with the work in this post aimed at impact.
The operational problem was overcommitment. Ai2 manages thousands of NVIDIA H100, B200, and B300 GPUs in clusters of 88 to 1024 units, serving roughly 150 internal researchers whose work spans large language and vision-language model training, robotics reinforcement learning simulation, and post-training for scientific agentic use cases. Outstanding requests at any moment seek 2-3x more GPUs than are available, meaning every available GPU hour has two or three workloads competing for it.
Two pathologies followed from the old design, in which workloads could opt out of preemptability. GPU squatting saw users park no-op workloads they could connect to when needed, because low-latency debug launches were not reliably available. Priority inflation followed, with eventually 100% of scheduled workloads running at HIGH priority, starving lower priority levels. Because preemptability was optional, on-call engineers spent a majority of ticket response time negotiating organized shutdowns of non-preemptable workloads on hosts with known maintenance problems.
Hugging Face Report Shows How Budgets Replace Scheduling Arguments
The team frames the failure as a tragedy of the commons: individuals competing over a scarce shared resource and, by maximizing individual outcomes, producing a non-optimal global result. It cites the 2011 Dominant Resource Fairness paper by Ghodsi and coauthors, which recounts a search company that granted dedicated machines only when users could guarantee high utilization, then found users sprinkling their code with infinite loops to inflate utilization levels.
Ai2's first correction was coarse, assigning teams monopolies over sets of GPUs. That produced idle hardware because research demand is seasonal: teams are ready to run experiments at different times, so a monopoly guarantees periods with no ready jobs while another team waits. Ai2 describes this as manually solving a knapsack problem, fitting dynamically changing research needs into a static schedule.
The replacement allocates GPU time rather than GPUs. The stated rationale is that predicting demand would require knowing the results of novel science experiments, while priority across research efforts is a strategy question that can be debated in advance. Leadership sets budgets, and managers proportionally allocate GPU time to the projects and researchers they oversee. Under the new system, nothing is free, so any trick to obtain GPU time draws from the benefiting user's allocation.
The fair-share algorithm itself is not presented as novel. Ai2 places it in a lineage running back to the 2009 Hadoop Fair Scheduler and notes the same approach is in active use in SLURM's Fair Tree and YARN's Fair Scheduler. What is new, per the post, is the input: a tree mirroring the research program structure, with weights set by managers rather than static quotas. The scheduler tracks occupancy over a sliding lookback window defaulting to 7 days and sorts workloads from under-utilized allocations above over-utilized ones. Allocation authority is tiered, with decisions made by a lead researcher within a project, a principal investigator within a program, and across programs by a lead program manager or the CEO.
Related: SEC Proposes Crypto Asset Rules for New Offering Exemptions
Hugging Face Report Quantifies the Time-Slicing Contract
The scheduling contract addresses long-running distributed training jobs, which Ai2 says regularly run for hours, days, and sometimes weeks. Once scheduled, such a workload could hold GPUs for a week or more with no opportunity for others to receive budgeted time, the same property that enabled squatting and forced on-call negotiation. Workloads now declare a minimum runtime, or the shortest occupancy needed to make meaningful progress, and are preemptible once that progress is banked. Setting minimum runtime to zero marks the time as unallocated and free but always preemptible. Resumable workloads are automatically re-queued.
The operational effects are concrete. Unhealthy hosts drain workloads as they reach minimum runtime, allowing repair activity to be automated; repairs requiring a human in the loop fell 74%. Debug workloads, which need few GPUs and under 15 minutes of protected runtime, saw p90 wait times fall from 2 hours to 30 seconds in production, versus a simulation prediction of 6 hours to 5 minutes from hand-crafted test scenarios. Ai2 notes the baseline's smaller debug sample carried higher variance. Median queue wait on the largest H100 cluster fell from 5 minutes to 24 seconds, and p90 wait fell about a third, from 2.8 hours to 1.8 hours. A researcher quoted in the post, Chris Clark, attributes a gain equivalent to roughly 30% more compute to reclaiming idle allocation and later bursting beyond allocation limits without preemption.
Against the three stated problems, Ai2 reports short debug workloads now start in under a minute, priority only sorts workloads within a team, and on-call toil dropped with automated draining. Priorities are not eliminated; their scope is narrowed.
Hugging Face Report Outlines Rollout Friction and Open Questions
The learning curve was steeper than assumed. Because rollout was incremental, researchers saw different behavior depending on which cluster they targeted, and interfaces retained old terminology such as workload priority whose meaning had changed. Documentation alone did not resolve the confusion. What worked, per the post, was live explanatory sessions using real examples, which shifted the team from a period of frustration and folk theories toward broader communication about experimental GPU needs. Post-launch visualizations showed how closely assigned GPU time tracked allocations and exposed the metric used to sort the queue.
For deeper context, see our AI Chips analysis: "Global Semiconductor Market Size, Share and Forecast Statistics by Country and Companies 2026-2030".
Not every use case improved. Interactive sessions for data analysis and testing training code could previously be held for up to a week; under time-slicing they are subject to the 8-hour protected runtime cap, after which a session becomes preemptible if it exceeds its allocation. Ai2 states it did not appreciate how dependent researchers were on the volatile state of these sessions, since preemption meant securing a new session and manually rebuilding state. After surveying researchers, it created two roadmap projects: a CPU-only cluster next to on-prem storage for data-prep dev sessions, and restorable sessions that can be preempted at the end of minimum runtime and restored elsewhere.
One unresolved risk is capacity fragmentation, which Ai2 says may raise queue wait times for the largest workloads. The stated intuition is that minimum runtime protection is being applied to jobs that previously relied on preemptible mechanisms to exceed team concurrent GPU limits. Those jobs could once be interrupted at any time, which wasted time but made it easier to schedule large jobs. The team says it is reproducing the problem in simulation while measuring production ground truth. The next stated aim is utilization: making bootstrapping, checkpointing, and training applications themselves as efficient as possible.
Hugging Face Signals Table
| Entity | Recent Focus | Geography | Source |
|---|---|---|---|
| Ai2 AI Infrastructure team | Replaced a priority-based GPU scheduler with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract | Not stated in source | Hugging Face |
| Ai2 research program | About 150 internal researchers spanning LLM and VLM training, robotics RL simulation, and scientific agentic post-training | Not stated in source | Hugging Face |
| GPU fleet | Thousands of NVIDIA H100, B200 and B300 GPUs in clusters of 88 to 1024 units | Not stated in source | Hugging Face |
| Fair-share scheduler lineage | Hadoop Fair Scheduler (2009), SLURM Fair Tree, YARN Fair Scheduler | Not stated in source | Hugging Face |
| Chris Clark | Researcher quoted on reclaiming roughly 30% more compute under the new scheduler | Not stated in source | Hugging Face |
The source does not specify locations for Ai2, its clusters, or its researchers, so no geography is asserted here.
Hugging Face Implementation Risks
The disclosure most relevant to anyone copying this design is that Ai2 reports its own results, and the article offers no independent verification of the 98% occupancy, 98% delivery, 74% reduction in human-in-the-loop repairs, or queue-latency figures. Several are self-measured against a baseline Ai2 itself describes as thin: the debug workload comparison rests on a small baseline sample the post flags for higher variance.
Additional coverage: Dell Q1 FY27 2026: $43.8B Quarter and $60B AI Server Guide
The unresolved technical risk is capacity fragmentation. Ai2 states that minimum runtime protection may be shielding the very jobs that previously absorbed frequent preemption, leaving the scheduler fewer opportunities to interrupt many jobs at once to place a large pending workload, which could increase wait times for the biggest jobs. That risk is under investigation rather than resolved.
A second risk is behavioral, not algorithmic. Ai2's stated strategy is to make gaming the scheduler more expensive than honestly arguing for a larger budget, and it acknowledges that users who lose a zero-sum exchange tend to look for new workarounds. The post also concedes it is constantly iterating on budget review, meaning the governance process, not the code, carries the load.
A third risk is change management. Incremental rollout, retained legacy terminology, and a steeper-than-assumed learning curve all surfaced as adoption friction that documentation did not solve. Rollout also began only at the end of July, so the 30-day test period is short relative to training workloads that can run for weeks.
Editorial independence disclosure: this article analyzes a publicly posted engineering account on Hugging Face; no relationship with Ai2, Hugging Face, or any named vendor influenced its preparation. Source: Hugging Face.
What This Means for Practitioners
For teams running shared GPU fleets, the transferable idea is not the fair-share algorithm, which Ai2 itself traces to Hadoop, SLURM, and YARN, but the funding precondition: nothing protected is free, so every request charges an owner's budget. That reframes scheduling disputes as budget decisions managers can make in advance, and it makes squatting and priority inflation self-defeating rather than policy violations to police. The second transferable idea is the minimum-runtime contract, which converts "how long might this run" into an explicit, resumable commitment and is what enabled 74% of repairs to proceed without a human. The caveat is that interactive sessions lose protected runtime and capacity fragmentation remains open.
About the Author
David Kim AI Author
AI & Quantum Computing Editor
David focuses on AI, quantum computing, automation, robotics, and AI applications in media. Expert in next-generation computing technologies.
David Kim is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →
Frequently Asked Questions
What did Ai2 change about its GPU scheduling?
Ai2's AI Infrastructure team replaced a priority-based scheduler with a system built on three components: GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract requiring each workload to declare a minimum runtime in exchange for protected occupancy. Under the new system, every protected request must be funded by an allocation or it is not shielded from preemption.
What hardware and how many researchers does Ai2's cluster serve?
Ai2 manages thousands of NVIDIA H100, B200 and B300 GPUs arranged in clusters ranging from 88 to 1024 GPUs, serving about 150 internal researchers whose work spans LLM and VLM training, robotics reinforcement learning simulation, and post-training for scientific agentic use cases. Demand exceeds supply, with outstanding requests seeking 2-3x more GPUs than are available.
What results does Ai2 report after the rollout?
Over a 30-day test period, teams received 98% of the GPU hours they were owed, with 13 of 15 team allocations at 95% or more and the worst case at 90%. Occupancy held steady at 98% before and after the change, and 18% of delivered GPU time was unallocated. Debug workload p90 queue wait time fell from 2 hours to 30 seconds in production, and repairs requiring a human in the loop fell 74%.
Did every workload benefit from the new scheduler?
No. Interactive sessions for data analysis and testing training code could previously be held for up to a week; under time-slicing they are subject to the 8-hour protected runtime cap, after which a session becomes preemptible if it exceeds its allocation. Ai2 created two roadmap projects in response: a CPU-only cluster for data-prep dev sessions and restorable sessions.
What open risks does Ai2 flag about the new system?
Ai2 says it is investigating capacity fragmentation, which may raise queue wait times for the largest workloads, because minimum runtime protection may be shielding jobs that previously absorbed frequent preemption. Ai2 also reports its own results, and the post offers no independent verification of the occupancy, delivery, repair or queue-latency figures.