Skip to content
Trending storyDeveloping

Ai2 replaces priority GPU scheduler with time-budget fair-share allocation

7 articles3 sourcessince Oct 9Last article 2h ago ·

Overview

AISummary of 7 articles

Ai2's AI Infrastructure team has replaced its priority-based GPU cluster scheduler with per-project GPU-time budgets, hierarchical fair-share allocation, and a time-slicing contract, so that decisions about each project's GPU time are made through a transparent budgeting process rather than case-by-case.

The clusters range from 88 to 1024 GPUs, including H100, B200, and B300 units, and serve about 150 internal researchers facing demand two to three times available capacity. Under the new system, managers assign budgets to programs and projects, and the scheduler moves work from underused allocations ahead in the queue.

Ai2 says the old scheduler let every workload eventually run at high priority, so researchers kept idle jobs running to hold GPUs, and that the new design removes that incentive. In its own claims, median queue wait on its largest H100 cluster fell from 5 minutes to 24 seconds, and over a 30-day test teams received 98% of the GPU hours they were owed, with occupancy at 98%. These results are Ai2's self-reported figures; no independent evaluation is reported.

Written by AI from the articles below · updated Oct 9, 12:01 PM ET

Check the sources:

Developments

6 developments

  1. Oct 9, 11:43 AM ET · 1 article
    Ai2 shares design decisions and rollout lessons behind its scheduler
    Ai2: Ai2 publishes design decisions and rollout lessons for its scheduler
  2. Oct 9, 11:43 AM ET · 1 article
    Ai2 reports 30-day test of new GPU scheduler delivering 98% of owed GPU hours
    Ai2: Ai2's new scheduler delivers 98% of owed GPU hours in 30-day test
  3. Oct 9, 11:43 AM ET · 1 article
    Ai2 describes a fair-share GPU scheduler that assigns research budgets
    Ai2: Ai2 describes a fair-share GPU scheduler for research budgets
  4. Oct 9, 11:43 AM ET · 1 article
    Ai2 shares engineering work behind its GPU scheduler and reports reduced queue waits
    Ai2: Ai2 shares engineering behind GPU scheduler that cut queue waits
  5. Oct 9, 11:43 AM ET · 1 article
    Ai2 replaces its GPU scheduler after idle jobs hoarded capacity
    Ai2: Ai2 replaces its GPU scheduler after idle jobs hoarded capacity
  6. Oct 9, 4:00 AM ET · 2 articles
    Ai2 describes replacing its GPU cluster priority scheduler with a budget-based fair-share scheduler
    Ai2 (Allen Institute for AI): Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. Ai2
    Ai2 publishes design decisions and rollout lessons for its scheduler

    AIAi2 has published a blog post on the design decisions behind its scheduler, along with rollout lessons and next priorities. The post links to The text provided does not include further details of the scheduler's features or results.

  2. Ai2
    Ai2 describes a fair-share GPU scheduler for research budgets

    AIAi2 says managers now assign GPU-time budgets to research programs and projects. Its fair-share scheduler compares recent usage with those budgets and moves work from underused allocations ahead in the queue.

    Image from @allen_ai's post
  3. Ai2
    Ai2's new scheduler delivers 98% of owed GPU hours in 30-day test

    AIAi2 reports that over a 30-day test of its new scheduler, teams received 98% of the GPU hours they were owed, based on actual demand. Cluster occupancy stayed at 98%, and spare capacity went to interruptible work without drawing down team budgets.

  4. Ai2
    Ai2 shares engineering behind GPU scheduler that cut queue waits

    AIAi2 published the engineering details of its new GPU scheduler, which allocates compute across research teams. On its largest H100 cluster, median queue wait fell from 5 minutes to 24 seconds.

    Video from @allen_ai's post
  5. Ai2
    Ai2 replaces its GPU scheduler after idle jobs hoarded capacity

    AIAi2 says its old scheduler made every scheduled workload eventually run at HIGH priority. Researchers kept idle jobs running to reserve GPUs for experiments, because the incentives rewarded holding capacity even with no active work.

  6. Hugging Face Blog
    Ai2 replaces priority scheduler with GPU time budgets for cluster allocation

    AIAi2's AI Infrastructure team replaced its priority-based GPU cluster scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change turns decisions about how much GPU time each research project receives into a transparent administrative budgeting process. Its clusters, which range from 88 to 1024 GPUs including H100, B200, and B300 units, serve about 150 researchers facing demand two to three times available capacity.

  7. Ai2 (Allen Institute for AI)
    Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

    AIAi2's AI Infrastructure team replaced its priority-based scheduler for GPU clusters with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change moved debates over how much GPU time each research project deserves from case-by-case operational decisions into a transparent budgeting process. The clusters range from 88 to 1024 GPUs across NVIDIA H100, B200, and B300 hardware, and serve about 150 internal researchers.

Heat trend

Not enough continuous observations to show a trend yet.