Where to find Python HPC modules

Python HPC (high-performance computing) modules live in three main places: your system's package manager, the Python Package Index (PyPI), and your institution's HPC cluster itself. The fastest route depends on what you're running and where.

If you're on a university or research cluster, the module is often already installed. Log in and type module avail to see what's loaded. Look for names like python-mpi4py, python-numpy, or python-scipy. These are pre-compiled for your hardware and won't require you to build anything.

If the module isn't there, check PyPI at pypi.org. Search for the module name — for example, "mpi4py" or "numba". Most HPC modules have a PyPI page that shows installation instructions, dependencies, and which Python versions they support. Read the requirements carefully: some HPC modules need a C compiler or MPI library already installed on your system.

Key Takeaways

  • On HPC clusters, use module avail and module load to find and activate pre-installed Python HPC modules instead of installing your own.
  • PyPI at pypi.org is the central repository where you can search for HPC modules, read their documentation, and see what dependencies they need.
  • Before installing an HPC module with pip, check whether your system has the required C compiler, MPI library, or CUDA toolkit that the module depends on.
  • Virtual environments keep HPC modules isolated from your system Python, which prevents version conflicts when you're working on multiple projects.
  • The module's GitHub repository or documentation site usually shows the exact build steps and system requirements for your operating system.

Using module load on HPC clusters

Most university and national lab clusters use a module system to manage software versions. When you log in, type module avail python to list all Python-related modules. You'll see output like python/3.9, python/3.11, or python-mpi4py/3.9.

Load the one you need with module load python/3.11. This adds the module to your environment and makes the Python executable available in your path. Check what loaded by typing module list. If you need a specific HPC module like mpi4py or numba, search for it the same way: module avail mpi4py. If it exists, load it with module load mpi4py.

The advantage of using pre-built modules is that they're compiled for your cluster's hardware and linked against the correct MPI library. Building them yourself often fails because of missing dependencies or compiler mismatches. Ask your cluster's support team if you can't find what you need — they may have it installed under a different name or in a restricted module category.

Installing from PyPI with pip

If the module isn't available through your cluster's module system, you can install it from PyPI using pip. First, create a virtual environment so your installation doesn't interfere with the system Python. Type python -m venv hpc_env to create a folder called hpc_env. Then activate it with source hpc_env/bin/activate on Linux or Mac, or hpc_env\Scripts\activate on Windows.

Once the virtual environment is active, type pip install mpi4py (or whatever module you need). Pip will download the module from PyPI and install it into your virtual environment. Check the installation with python -c "import mpi4py" — if it runs without error, the module is ready.

Some HPC modules require a C compiler or system library before pip can build them. If installation fails, read the error message carefully. It usually says what's missing — for example, "mpi.h not found" means you need to install the MPI development headers. On Ubuntu, that's apt install libopenmpi-dev. On a cluster, it's usually module load openmpi before you run pip.

Checking dependencies before installation

Before you install an HPC module, visit its PyPI page or GitHub repository and read the "Requirements" or "Installation" section. HPC modules often depend on libraries that aren't installed by default. Common ones include a C or Fortran compiler, an MPI implementation like OpenMPI or MPICH, CUDA for GPU computing, or linear algebra libraries like BLAS and LAPACK.

On a cluster, these are usually available as modules. For example, if you're installing a module that needs CUDA, type module avail cuda first. Load the version you need with module load cuda/11.8. Then run pip install. The module system sets environment variables that tell pip where to find the libraries it needs.

On your own machine, you may need to install dependencies through your system's package manager. On macOS with Homebrew, that might be brew install open-mpi. On Ubuntu, it's apt install libopenmpi-dev. Check the module's documentation for the exact command for your operating system.

Finding modules on GitHub and documentation sites

The most detailed information about an HPC module usually lives on its GitHub repository or official documentation site, not on PyPI. Search for the module name plus "github" — for example, "mpi4py github". The repository's README file shows build instructions, system requirements, and examples of how to use the module.

Some HPC modules have their own documentation site. Numba, for instance, has a full guide at numba.readthedocs.io that explains how to install it on different systems and how to write code that uses it. These sites often have troubleshooting sections that cover common installation problems.

If you hit an error during installation, search the repository's issues page. Type your error message into the search box — someone else has probably hit the same problem and posted a solution. If not, you can open a new issue and the maintainers may help you diagnose it.

Working with conda instead of pip

Conda is an alternative to pip that often handles HPC module dependencies better, especially on your own machine. If you have Anaconda or Miniconda installed, you can search for modules on anaconda.org. Type conda search mpi4py to see what versions are available. Install with conda install mpi4py.

Conda pre-builds modules for common platforms and links them against compatible libraries, so you usually don't need a C compiler or system libraries installed separately. This makes it faster and more reliable than pip for complex HPC modules. Create a conda environment for your project with conda create -n hpc_project python=3.11 mpi4py numpy scipy, then activate it with conda activate hpc_project.

On HPC clusters, conda is less common because clusters usually have their own module system. But if your cluster has conda available, it can be a good option if the module you need isn't in the cluster's module system and pip installation is failing.

Troubleshooting installation failures

If pip or conda installation fails, the error message usually points to the problem. "Command 'gcc' not found" means you need a C compiler — install it with your system's package manager or load it as a module. "mpi.h not found" means the MPI development headers aren't installed — load the MPI module or install the dev package.

Version mismatches are common. If you're installing a module for Python 3.11 but your system has Python 3.9, pip may refuse to install it or install a version that doesn't work. Check which Python version is active with python --version. If it's wrong, load the correct Python module or activate a virtual environment that uses the right version.

Some HPC modules are only available as pre-built wheels for certain platforms. If you're on an unusual system or architecture, pip may not find a compatible version. In that case, you may need to build the module from source. Download the source code from GitHub, follow the build instructions in the README, and install it with pip install . from the source directory.

Frequently Asked Questions

What's the difference between module load and pip install?

Module load activates a pre-built version that's already on your cluster, while pip downloads and installs from PyPI. Module load is faster and avoids build errors, but only works if the module is already installed on your system. Pip works anywhere but requires dependencies to be installed first.

Can I use both module load and pip in the same project?

Yes. Load system modules first with module load, then create a virtual environment and use pip to install additional modules. The system modules set environment variables that pip can use to find libraries. This is the standard workflow on HPC clusters.

Why does pip say "no matching distribution found"?

The module may not be built for your Python version or operating system. Check the PyPI page to see which versions are available. If your Python version isn't listed, you may need to use a different version or build the module from source. On clusters, ask your support team if they have a pre-built version.

How do I know if an HPC module is installed correctly?

Type python -c "import modulename" and press Enter. If nothing happens, it's installed. If you get an ImportError, it's not. You can also type python -c "import modulename; print(modulename.__version__)" to see which version is installed.

Should I install HPC modules in a virtual environment?

Yes, on your own machine. Virtual environments prevent version conflicts and keep your system Python clean. On HPC clusters, it's optional but still recommended if you're working on multiple projects with different module versions.