Compute time on JSC hardware¶
Compute Time Projects¶
On this page, we give a loose overview over options for obtaining computational resources for the JSC systems. It is by no means legally binding, but should summarize key aspects for AI users.
Access to the system at JSC is only possible with an allocation of computational resources. Such allocations are called compute time projects. These compute time projects are allocated with different procedures. For different allocation types, different classes of users are eligible. Below, you will find a list, who can apply for which types of projects, and the approximate sizes.
The volume of compute time projects are measured in core-hours (CoreH). Unfortunately, this unit is not well comparable, as for example the number of GPUs might be the more decisive factor for computational performance and cost of the hardware than the number of cores. However, it remains the smallest common denominator. For practical purposes, it is helpful to always keep a suitable conversion factor, for example between CoreH and GPU-hours in mind. Below we provide a table with such comparisons for all JSC systems.
Types of Compute Time Projects¶
Here we give an informal orientation, for interested scientists. This is not official, but intended to give a coarse orientation. The official JSC page regarding compute time is found here.
Type: Test Projects
System: JURECA-DC CPU and GPU, JUWELS Cluster and Booster, JUSUF
Accelerator: Diverse
Eligible: Everybody
Size: Extremely small. Hundreds of Node-hours.
Deadlines: Application possible at any time.
Proposal: Short project description
More Information: https://www.fz-juelich.de/en/ias/jsc/systems/supercomputers/ call-for-applications-for-test-projects-with-jsc-supercomputing-and-support-resources
Type: VSR
System: JURECA-DC CPU and GPU
Accelerator: NVIDIA A100 GPU 40 GB
Eligible: Scientists from FZJ
Size: No official limits, a good orientation is 10.000 GPUh – 150.000 GPUh.
Deadlines: Two deadlines per year, mid-February, and mid August. Projects start May and Nov 1st.
Proposal: Scientific proposal including resource estimation
More Information: https://www.fz-juelich.de/en/ias/jsc/systems/supercomputers/apply-for-computing-time/vsr
Type: GCS (Gauss Centre for Supercomputing)
System: JUWELS Booster (JUWELS Cluster can be interesting for data processing)
Eligible: Employees of German research institutions
Size: No official limits, a good orientation is 20.000 CoreH – 600.000 GPUh for regular projects, and 600.000 - 1.200.000 GPUh for large projects (JUWELS Booster).
Deadlines: Two deadlines per year, mid-February, and mid August. Projects start May and Nov 1st.
Proposal: Scientific proposal including resource estimation
More Information: https://www.fz-juelich.de/en/ias/jsc/systems/supercomputers/apply-for-computing-time/gcs-nic
Type: HAICORE (Helmholtz AI Compute Resources)
System: JUWELS Booster
Accelerator: NVIDIA A100 GPU 40 GB
Eligible: Employees of Helmholtz research institutions
Size: 5.000 GPUh
Deadlines: Application possible any time.
Proposal: Project summary
More Information: Lightweight Projects at https://www.helmholtz.ai/themenmenue/you-helmholtz-ai/computing-resources/index.html
Type: WestAI
System: JURECA-WestAI
Accelerator: NVIDIA H100 GPU 94 GB
Eligible: Everybody
Size: 10.000 GPUh
Deadlines: Applications not yet possible.
Proposal: Project summary
More Information: https://westai.de/
The review process¶
In a full project proposal, it is required to convince reviewers, that your project is well thought through and is worthy of receiving a publicly funded, shared resource. Loosely speaking, you must bring across three points.
- You have a exciting scientific case
- You can make good use of the infrastructure
- You have a plausible plan how much resources will be required.
The first case of scientific excellence is typically not very difficult. In many cases, review processes of scientific projects have already shown this is set. For the second point, it is advised to show a proof of the computational performance of your code(s) in your application. Canonically, this is done as a scaling plot, where the throughput is plotted as a function of the number of GPUs or nodes. The third point is notoriously difficult, especially as some experiments depend on the outcome of other experiments. Keep in mind that your plan should resemble your best guess, and that deviations are absolutely possible.
FLOPs, Core-hours and GPU-hours¶
The default units, in which computational resources are measured, are core-hours, or CoreH. In a time before GPUs and other accelerators were standard, it was a reasonable, and reasonably comparable metric for computational resources. With the number of (physical) cores, this unit is related to the amount of usage time of an entire compute node, or the NodeH. At the point of writing this document, a common unit of computational resources is one hour of NVIDIA A100 GPUs. The number of GPUh, of course, related to the number of NodeH by the number of GPUs per node, and to the CoreH by the ratio of cores per GPU.
Beyond the usage of device-related metrics, possible metrics are related to the number of Floating Point Operations. However, these numbers are also not ideally comparable, they are computed assuming ideal usage of each device in 64-bit floating point precision. Another common unit of compute is PetaFLOP/s-days, a unit proposed for example by OpenAI.
The following table, we list conversion factors from Node-hours all other units, where available
| System | Node | Accelerator | NodeH | CoreH | GPUh | EFLOP | PFLOP/s d |
|---|---|---|---|---|---|---|---|
| JUWELS | Booster | NVIDIA A100 40 GB | 1 | 48 | 4 | ||
| JUWELS | Cluster GPU | NVIDIA V100 16 GB | 1 | 40 | 4 | ||
| JURECA-DC | GPU | NVIDIA A100 40 GB | 1 | 128 | 4 | 0.30 | |
| JURECA-DC | WestAI | NVIDIA H100 94 GB | 1 | 32 | 4 | ||
| JUSUF | GPU | NVIDIA V100 | 1 | 128 | 1 |