AI.BAY - Bayerische KI-Basismodell-Initiative#
Getting access to your AI.BAY project#
Invite users#
To get access, the PI of a project must first log in to our HPC portal at https://portal.hpc.fau.de. They are then assigned to the corresponding project. Once this is done, the Manager tab becomes visible in the portal and the PI can invite users to the project. Further information about the HPC portal can be found here.
When inviting a user, the PI or manager of a project must work through a decision tree to ensure that export control requirements are met.
Login#
You log in to NHR@FAU systems via SSH. A short guide to creating an SSH key and setting up your SSH configuration can be found here.
Resources#
GPUs#
The H200 partition of Helma is currently dedicated to the Bayerische KI-Basismodell-Initiative. It comprises 96 nodes, each equipped with four H200 GPUs (141 GB HBM3e).
We expect a further 1,000 B200 GPUs to become available by the end of 2026.
Storage#
In addition to your home directory $HOME and $HPCVAULT (see filesystems), 4.3 PB of NVMe storage is available. You can use this storage via workspaces.
We currently use an inode quota (limit of files and folders) and a volume quota on /hnvme.
As an AI user you might be dealing with datasets containing huge amounts of files. Please have a look at our documentation on that topic to avoid problems.
Each user also has a $WORK directory with a group/project quota of 3 TB.
Compute time quotas#
For the Health and the Robotics & Perception clusters, a single monthly quota applies to all projects combined. All other projects have an individual rolling monthly quota, consisting of the granted allocation plus any GPU hours left over from the previous month. Unused hours are carried over for one month only; they do not accumulate over several months.
Once the quota is reached, users can only submit jobs to the preempt queue (see below and Slurm). An (QOSGrpGRESMinutes) in the output of squeue indicates that the quota was hit.
Quotas have been disabled for the initial phase of the project.
Queuing system#
The maximum number of running and pending jobs per user is 1,000.
Once your project's quota is exhausted, you can still submit to the preempt queue:
Please note that jobs in this queue can allocate one node only and may be terminated after two hours of run time in favor of regular jobs.
Staying up to date#
Information on down times is published on the NHR@FAU website. Please also make sure you subscribe to the mailing lists nhr-users and nhr-sysannounce in the HPC portal under Profile. We use these lists to share important information with our HPC users from time to time.
Events#
We would like to draw your attention to the HPC-Cafe offered by NHR@FAU on a regular basis. Especially, Beginner’s Introduction "HPC in a nutshell" might be of interest for new users. All relevant information can be found here.