Jump to HeaderJump to Main ContentJump to Footer
Michigan State University
Michigan State University
ICER Institute for Cyber-Enabled Research
  • HPCC Status
  • ICER Docs

ICER Institute for Cyber-Enabled Research

  • About
    • Overview
    • News
    • Affiliations
    • Employment Opportunities
    • FAQ
    • Newsletters
    • 20th Anniversary
    Supercomputer hardware with blue and green lights and many wires connect it together.

    View and download photos of ICER's hardware.

    Photo Gallery
  • |
  • System Status
    • Overview
    • HPCC Service Status
    • ICER Status Dashboard
  • |
  • For HPCC Users
    • Overview
    • Getting Started
    • HPCC System Info
    • HPCC User Documentation
    • Buy-In Options
    • Topic of the Month
  • |
  • Research Services
    • Overview
    • Research Highlights
    • Publications
    • Citing ICER
    • Grant and Research Assistance
    • Academic Research Consulting Services
    • Opportunities and Communities
  • |
  • Training and Education
    • Overview
    • Webinars and Seminars
    • Asynchronous Tutorials
    • Classroom Support
  • |
  • Contact
    • Overview
    • Directory
ICER Institute for Cyber-Enabled Research
  • About
  • System Status
  • For HPCC Users
  • Research Services
  • Training and Education
  • Contact

< About

  • Overview
  • News
  • Affiliations
  • Employment Opportunities
  • FAQ
  • Newsletters
  • 20th Anniversary
Supercomputer hardware with blue and green lights and many wires connect it together.

View and download photos of ICER's hardware.

Photo Gallery

< System Status

  • Overview
  • HPCC Service Status
  • ICER Status Dashboard

< For HPCC Users

  • Overview
  • Getting Started
  • HPCC System Info
  • HPCC User Documentation
  • Buy-In Options
  • Topic of the Month

< Research Services

  • Overview
  • Research Highlights
  • Publications
  • Citing ICER
  • Grant and Research Assistance
  • Academic Research Consulting Services
  • Opportunities and Communities

< Training and Education

  • Overview
  • Webinars and Seminars
  • Asynchronous Tutorials
  • Classroom Support

< Contact

  • Overview
  • Directory
  • HPCC Status
  • ICER Docs
ICER > For HPCC Users > Topic of the Month >

Checklist to Improve Your Jobs Scheduling Time

Checklist to Improve Your Job's Scheduling Time

Here is a checklist to help improve your job’s scheduling time:

1. Run command “sq <NetID>” or “squeue -u <NetID>” to show the jobs in the queue. If a job is Pending, the last column will show the reason provided by SLURM. Here are common reasons with their meaning:

  • a. Priority: Job queued behind some higher priority jobs. For information on how job priority is computed, review Job Priority Factors.
  • b. Resources: Job waiting for available resources. Users can use the powertools command “node_status” to see a current list of available node status. For an overview of resources, see Cluster Resources.
  • c. Dependency: Job waiting for the jobs it depends on to complete.
  • d. JobHeldAdmin: Job held by a system administrator. Contact system administrator for assistance using the Contact Forms. 
  • e. JobHeldUser: Job held by you. Run “scontrol release <job_id>” to release the hold.
  • f. JobArrayTaskLimit: You reached the job array task limit per fairshare policy. Job waiting for array jobs to complete.
  • g. QOSMaxCpuPerUserLimit: You reached the maximum CPU per user limit per fairshare policy.
  • h. QOSMaxJobsPerUserLimit: You reached maximum jobs per user limit per fairshare policy.
  • i. BadConstraints: Job constraints cannot be satisfied. You can put any constraints for the particular type of nodes, number of nodes, number of CPUs per node, etc. to control the way your jobs are executed. Check the job script to ensure the requested resources are available.

2. Job submission during peak usage times may contribute to longer wait times. Check the ICER Dashboard to see the current system load. Compare your expected queue time to the “Queue Times” data, which provides the average queue time of the last 100 completed jobs of similar type, and adjust your expectation of job queue time accordingly.

3. Review the HPCC fairshare policy. If you find that your jobs are waiting in the queue longer than another user’s similar jobs, it may be due to your large resource usage recently. Users can find the FAIRSHARE contribution to a job priority by running command "sprio -u $USER" and compare the fairshare numbers.

4. Here are some tricks that may lead jobs to be scheduled faster:

  • a. Request a <4hrs job time. Short jobs can be scheduled on buyin nodes as available.
  • b. Avoid requesting high-demand resources. High demand for certain resources, like GPUs, increases queue times due to availability limitations. Users can balance a job’s queue and execution times between running with and without a GPU for the optimal situation.
  • c. Avoid overestimating needed resources. A more accurate estimation of needed resources (i.e. time, CPUs and GPUs number, GB of memory) can reduce queue times since larger size jobs are still more difficult to be scheduled due to the availability of the resources.
  • d. Avoid unnecessary constraints in the job script. Fewer constraints will give SLURM flexibility to collect the available resources for the jobs leading to a faster schedule.

Xiaoge Wang
Research Consultant
ICER
 

Michigan State University
  • Documentation Homepage
  • HPCC Service Status
  • ICER Newsletter

Contact us

Contact Form Link

Address

Biomedical & Physical Sciences Building 567 Wilson Road, Room 1440 East Lansing, Michigan 48824-1226

Follow Us

  • Visit our Facebook page
  • Visit our page on X
  • Visit our Instagram page
  • Visit our LinkedIn page
  • Visit our YouTube page

If you're having accessibility issues, please let us know.

Michigan State University
  • Contact Information|
  • Site Map|
  • Privacy Statement|
  • Site Accessibility|
  • Call MSU: (517) 355-1855|
  • Visit: msu.edu|
  • Notice of Nondiscrimination|

SPARTANS WILL|© Michigan State University|