Machine learning is a subfield of artificial intelligence that uses algorithms trained on data sets to create models that enable machines to perform tasks that would otherwise only be possible for humans. These tasks can include categorizing images, analysing data, or predicting price fluctuations.
The eResearch team manages access to several private VMs and a dedicated machine that is suitable for machine learning tasks.
For anyone new to machine learning, we recommend some starting resources to learn key concepts and the technologies involved: the Machine Learning Crash Course by Google and Kaggle Data Science Education.
GPUs on the RCH
The Research Compute Hub has a number of GPUs that researchers can use for Machine Learning. For more information on how the GPUs are setup on the RCH, please refer GPUs on the RCH page.
SLURM and Scheduling
Since it is a shared resource, we need a way to ensure fair access amongst the users. With a few users some kind of board or mailing list to request a turn on the machine is fine. With an increasing amount of users, we need a formal workload manager providing a submission queue. Instead of running jobs directly on the GPUs, users need to submit the job to a workload manager which will manage who runs when. A time limit (usually referred as "wallclock time") is also enforced so people do not wait too long in the queue. Use of a job scheduler will also
- enable us to figure out how busy is the machine
- use the machine more efficiently as no one will have to figure out if the machine is busy or wait for a message that it is their turn
The selected workload manager is slurm. Slurm is open source and an industry standard in the HPC world. People using NeSI will be familiar with it, in turn users of the RCH will become familiar with the technology used at NeSI and many other facilities (according to wikipedia slurm is used by about 60% of the TOP500 supercomputers).
Job Submission with Slurm
To submit a job, you need to prepare a small text file which contains a script to run and slurm instruction describing the job and its requirements and then putting it in the queue with the appropriate command. A good summary of slurm commands and options can be found on the UC RCH Wiki, other pages with interesting examples can be found at compute Canada and university of Cambridge.
Public Services
Public services normally have an associated cost that is not covered by the University. University services are free of charge but may have constraints on the type and quantity of hardware available as well as the duration of your projects.
- Google Colaboratory: Colaboratory allows you to write and execute Python in your browser with free access to GPUs and easy sharing. With Colab you can harness the full power of popular Python libraries to analyse and visualize data.
- Amazon Web Services: Amazon Web Services offers a broad set of machine learning services and supporting cloud infrastructure.
- Microsoft Azure Machine Learning Studio: The Azure Machine Learning service empowers developers and data scientists with a wide range of productive experiences for building, training, and deploying machine learning models.
- Google Cloud Machine Learning: Google Cloud offers AI and machine learning products for developers, data scientists, and data engineers.
- IBM Watson: Watson is IBM’s portfolio of enterprise-ready pre-built applications, tools, and runtimes. With Watson you can infuse AI into your applications to make predictions or automate decisions and processes.
- NeSI: NeSI provides a national platform of shared high performance computing tools and eResearch services. They further have many resources dedicated to machine learning.
This list does not contain information on Generative AI tools at UC which can be found here.