HPC Cluster

HPC: 

Welcome to the high performance computing (HPC) community site at CCHMC! We operate as the Research Computing group under Information Systems for Research (IS4R).

Currently, we maintain one Red Hat (RHEL 9) Linux based HPC cluster and one AI/ML GPU cluster for research. Our primary HPC cluster environment currently has 2500+ cores and is heterogeneous with both large-memory SMPs totaling 35TB of RAM across 80 nodes. The primary connection for the cluster nodes is a high speed Ethernet with 10-25Gbps and the scheduler / resource manager is IBM LSF. This environment also contains nodes with GPU (NVIDIA) capabilities. The cluster contains 10 x V100 with 32GB of RAM, 20 x A100 with 40GB/80GB RAM, and 10 x  RTX 6000 Blackwell with 96GB RAM. The AI/ML cluster is tailored for containerized GPU workloads and currently has 8 x H200 GPUs, 32 x H100 GPUs and 4x V100 GPUs.

The software available on the cluster is installed upon request and is managed via TCL modules, conda environments, and containers. We have several versions of R/Rstudio, Python, Nextflow, Picard, and Samtools as examples. Also, there is a Web interface "HPC OnDemand" where most tools and desktops can be deployed from any web browser with easy use. The AI/ML cluster utilizes Run:AI on top of Kubernetes and features a rich web interface along with CLI and API based access. We also have an LLM platform built on Ollama and vLLM and leveraging Run:AI - this allows us to serve many open models at scale.

Cluster storage is NFS (Network File System) based where each user has 100GB allocated for home and 5TB for scratch. Also, data shares can be requested for additional storage and shared with multiple users and other institutions using Globus or Active MFT.

The clusters are open to all CCHMC employees and collaborators with a valid use case.

If you have a question that is unanswered after reviewing this information, please email our support system at help-cluster@bmi.cchmc.org.