Datasets for training machine learning algorithms¶
Important
The data can only be used for non-commercial, scientific or educational purposes.
Please pay close attention to the terms of use and the licensing information included for each dataset. Occasionally, citation to the original work is required.
By downloading or accessing the data available centrally, you agree to the terms and conditions published by the copyright holder of the data. The responsibility is on the users to check the terms of use of the data and to make sure their use case is aligned with what is permitted by the copyright owner. In case explicit terms of use of the data is missing from the provider's side, the NonCommercial-ShareAlike (NC-SA) attribute of the Creative Commons 3.0 license applies.
Datasets are provided through modules and by loading a module environment variables are populated pointing to the dataset. It is important that you load modules to use the datasets and don't use the path directly as:
- The path to datasets are subject to change, by using the environment variables from the modules your code will continue to run.
- By loading the modules we get an idea of what datasets are in use.
- Notice about changes to specific datasets will be in connection to loading the module.
If you think a required, publicly available dataset is missing please contact us through the support form in SUPR. We will evaluate if it is something that should be added.
Dataloading pipeline¶
The dataloading pipeline is a crucial component for computational efficiency in supervised training. Some recommendations for clusters are (in order of usefulness):
- Reading directly from Mimer is usually fast enough, that way you'll save
time that would have been used to transfer to
$TMPDIRor/dev/shm. - Many small files are slow prefer to bunch the data into only a few larger
files (WebDataset, HDF5, zip, Arrow (e.g.
datasetspackage), TFRecord). - Use compression wisely:
- Jpeg, png, gz, or other files that don't benefit from (further)
compression shouldn't be compressed (
zip -0). - Very compressible formats (csv, mmcif, txt, ...) can potentially benefit from compression.
- Jpeg, png, gz, or other files that don't benefit from (further)
compression shouldn't be compressed (
- When shuffling data zip, is preferred over tar and if you can, shuffle before processing data points.
- Using a profiler is very useful to get an idea what works or not and what part of the pipeline it is that takes time (e.g. see TensorBoard).
- Don't spend more time optimising the dataloading pipeline than you'll save from the optimisation.
- Batch size, number of preprocessing workers and prefetching strategy can
have significant impact as a rule of thumb:
- If loading data is very fast, use the 0th (current thread) as worker.
- Otherwise you can spend some effort in setting up prefetching and loading multiple workers.
Mounting compressed datasets (ZIP)¶
Some datasets are distributed as .zip archives for efficient storage and transfer.
On Vera, you can access them directly without extracting, either:
- on the host using the built-in
fuse-archivecommand, or - inside Apptainer containers using the
--fusemountoption.
See more information in the filesystem documentation.
Available datasets¶
The following datasets are currently available:
AlphaFold3 data¶
AlphaFold3-data/3.0
These are the datasets needed to run inference with AlphaFold3, note that we are not allowed to distribute the model weights and they thus are not included.
The following databases have been: (1) mirrored by Google DeepMind; and (2) in part, included with the inference code package for testing purposes, and are available with reference to the following:
- BFD (modified), by Steinegger M. and Söding J., modified by Google DeepMind, available under a Creative Commons Attribution 4.0 International License. See the Methods section of the AlphaFold proteome paper for details.
- PDB (unmodified), by H.M. Berman et al., available free of all copyright restrictions and made fully and freely available for both non-commercial and commercial use under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication.
- MGnify: v2022_05 (unmodified), by Mitchell AL et al., available free of all copyright restrictions and made fully and freely available for both non-commercial and commercial use under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication.
- UniProt: 2021_04 (unmodified), by The UniProt Consortium, available under a Creative Commons Attribution 4.0 International License.
- UniRef90: 2022_05 (unmodified) by The UniProt Consortium, available under a Creative Commons Attribution 4.0 International License.
- NT: 2023_02_23 (modified) See the Supplementary Information of the AlphaFold 3 paper for details.
- RFam: 14_4 (modified), by I. Kalvari et al., available free of all copyright restrictions and made fully and freely available for both non-commercial and commercial use under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication. See the Supplementary Information of the AlphaFold 3 paper for details.
- RNACentral: 21_0 (modified), by The RNAcentral Consortium available free of all copyright restrictions and made fully and freely available for both non-commercial and commercial use under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication. See the Supplementary Information of the AlphaFold 3 paper for details.
CO3Dv2 data¶
CO3Dv2-data/20231130-zip: Provided as a number of uncompressed zipfiles.
Common Objects in 3D version 2. Licensed under CC BY-NC 4.0. If you use the dataset please use the following citation:
@inproceedings{reizenstein21co3d,
Author = {Reizenstein, Jeremy and Shapovalov, Roman and Henzler, Philipp and Sbordone, Luca and Labatut, Patrick and Novotny, David},
Booktitle = {International Conference on Computer Vision},
Title = {Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction},
Year = {2021},
}
ImageNet 1k data¶
Gated dataset
Access to this dataset is gated and not provided openly to all users. To get access join the related license group in SUPR: https://supr.naiss.se/group/215/member/subscribe/?.
Please cite the main ImageNet paper (Deng et al. CVPR 2009) if you use ImageNet data, in addition to specific follow-up papers:
@inproceedings{deng2009imagenet,
title={Imagenet: A large-scale hierarchical image database},
author={Deng, Jia and Dong, Wei and Socher, Richard and Li, Li-Jia and Li, Kai and Fei-Fei, Li},
booktitle={2009 IEEE conference on computer vision and pattern recognition},
pages={248--255},
year={2009},
organization={IEEE}
}
ImageNet-1k-data/20250917-hf¶
Dataset based on ILSVRC/imagenet-1k on HuggingFace hub. To use:
nuScenes data¶
Gated dataset
Access to this dataset is gated and not provided openly to all users. To get access join the related license group in SUPR: https://supr.naiss.se/group/220/member/subscribe/?.
Access to this dataset is covered by a license group in SUPR.
nuScenes-data/1.0-map-1.3-zip: Main data as uncompressed zip files with map extension.nuScenes-lidarseg-data/1.0-zip: Lidarseg extension as uncompressed zip files.nuScenes-panoptic-data/1.0-zip: Panoptic extension as uncompressed zip files.
Non-commercial use only, for detailed terms-of-use see: https://www.nuscenes.org/terms-of-use
- Overview: https://www.nuscenes.org/overview
- Trainval (700+150 scenes) is packaged into 10 different archives that each contain 85 scenes.
- Test (150 scenes) is used for challenges and does not come with object annotations.
- Mini (10 scenes) is a subset of trainval used to explore the data without having to download the entire dataset.
- The meta data is provided separately and includes the annotations, ego vehicle poses, calibration, maps and log information.
- Citation: please use the following citation when referencing nuScenes:
@article{nuscenes2019,
title={nuScenes: A multimodal dataset for autonomous driving},
author={Holger Caesar and Varun Bankiti and Alex H. Lang and Sourabh Vora and
Venice Erin Liong and Qiang Xu and Anush Krishnan and Yu Pan and
Giancarlo Baldan and Oscar Beijbom},
journal={arXiv preprint arXiv:1903.11027},
year={2019}
}
Deprecated datasets¶
The following datasets are deprecated and are planned to be removed. If you are using any of these and want them to remain, contact support.
Argoverse2-data/deprecatedBDD100K-data/deprecatedBIOSCAN-5M-data/20260109-hf-deprecatedBigEarthNet-S2-data/deprecatedCryoPPP-data/deprecatedE-GMD-data/1.0.0-deprecatedKITTI-data/deprecatedLSUN-data/deprecatedLyft-Level5-data/deprecatedMAN-TruckScenes-data/deprecatedMPI-Sintel-data/deprecatedMegaDepth-data/deprecatedMegaScenes-data/deprecatedMicrosoft-COCO-data/deprecatedObject365-data/deprecatedOpenImages-V6-data/deprecatedOpenSLR-SLR12-data/deprecatedPandaSet-data/deprecatedPhysioNet-challenge-2018-data/1.0.0-deprecatedPlaces365-data/deprecatedScene-Flow-data/mix-deprecatedSlakh2100-redux-data/deprecatedWaymo-Open-Dataset-data/1.2.0-deprecatedYouTube-VOS-data/2019-deprecatedZenseact-Open-Dataset-data/deprecated
Available LLM model weights¶
These models provide the path to the model weights and can be used with the
transformers package or vLLM directly. Each of these modules provides
environment variables:
HF_NAME: Name of the model e.g.openai/gpt-oss-120bHF_MODEL: Path to the modle weights
as well as possibly a few more specific ones related to the model name. To see
exactly what a module does use module show <specify module here>, e.g.
module show openai-gpt-oss-120b-model/20250826-hf.
Available model modules:
openai-gpt-oss-120b-model/20250826-hf
Deprecated LLM model weights¶
The following models are deprecated and are planned to be removed. If you are using any of these and want them to remain, contact support.
HuggingFaceTB-SmolLM2-135M-Instruct-model/12fd25f-deprecatedHuggingFaceTB-SmolLM2-135M-model/93efa2f-deprecatedHuggingFaceTB-SmolLM3-3B-Base-model/d78a42f-deprecatedHuggingFaceTB-SmolLM3-3B-model/a07cc9a-deprecatedQuantTrio-GLM-4.5-Air-GPTQ-Int4-Int8Mix-model/1c98fb7-deprecatedQwen-Qwen2-0.5B-model/91d2aff-deprecatedQwen-Qwen2.5-0.5B-Instruct-model/7ae5576-deprecatedQwen-Qwen2.5-VL-3B-Instruct-model/6628554-deprecatedQwen-Qwen3-0.6B-Base-model/da87bfb-deprecatedQwen-Qwen3-Coder-30B-A3B-Instruct-model/573fa39-deprecatedQwen-Qwen3-VL-235B-A22B-Instruct-model/a85c315-deprecatedQwen-Qwen3.5-122B-A10B-GPTQ-Int4-model/5b9f005-deprecatedQwen-Qwen3.5-27B-FP8-model/2e1b213-deprecatedQwen-Qwen3.5-27B-GPTQ-Int4-model/507bda6-deprecatedQwen-Qwen3.5-27B-model/b7ca741-deprecatedQwen-Qwen3.5-397B-A17B-GPTQ-Int4-model/b54fd48-deprecatedQwen-Qwen3.5-9B-model/c202236-deprecatedQwen-Qwen3.6-27B-model/6a9e13b-deprecatedQwen-Qwen3.6-35B-A3B-model/995ad96-deprecatedRedHatAI-Qwen3.6-35B-A3B-NVFP4-model/e850c69-deprecatedRedHatAI-gemma-4-26B-A4B-it-NVFP4-model/95d1e59-deprecatedRedHatAI-gemma-4-31B-it-NVFP4-model/c490598-deprecatedSnowflake-snowflake-arctic-embed-l-v2.0-model/ac6544c-deprecatedcyankiwi-Qwen3.6-27B-AWQ-INT4-model/c9b937c-deprecateddeepseek-ai-DeepSeek-V3.2-Exp-model/194c67e-deprecatedllava-hf-llava-v1.6-vicuna-13b-hf-model/2c9a8b5-deprecatedllava-hf-llava-v1.6-vicuna-7b-hf-model/c9f1252-deprecatedlmsys-vicuna-7b-v1.3-model/236eeea-deprecatedmistralai-Devstral-Small-2-24B-Instruct-2512-model/1da725e-deprecatedmoonshotai-Kimi-K2-Thinking-model/6126819-deprecatedneuralmagic-Llama-3.2-11B-Vision-Instruct-quantized.w4a16-model/7f66874-deprecatedneuralmagic-Llama-3.3-70B-Instruct-quantized.w8a8-model/dc36722-deprecatedneuralmagic-Meta-Llama-3.1-405B-Instruct-quantized.w4a16-model/9db6306-deprecatedneuralmagic-Meta-Llama-3.1-70B-Instruct-quantized.w4a16-model/b5b7bd4-deprecatedneuralmagic-Meta-Llama-3.1-8B-Instruct-quantized.w4a16-model/a7c0994-deprecatednomic-ai-nomic-embed-text-v1.5-model/e5cf08a-deprecatednomic-ai-nomic-embed-text-v2-moe-model/1066b65-deprecatedopenai-community-gpt2-model/607a30d-deprecatedopenai-gpt-oss-120b-model/b5c939d-deprecatedopenai-gpt-oss-120b-model/bc75b44-deprecatedopenai-gpt-oss-20b-model/cbf31f6-deprecatedopenai-gpt-oss-20b-model/6cee5e8-deprecatedsentence-transformers-all-MiniLM-L6-v2-model/c9745ed-deprecatedunsloth-Llama-3.2-11B-Vision-Instruct-model/677b0c1-deprecatedunsloth-Llama-3.2-1B-Instruct-bnb-4bit-model/117a0a1-deprecatedunsloth-Llama-3.2-1B-Instruct-model/d2b9e36-deprecatedunsloth-Llama-3.2-3B-Instruct-bnb-4bit-model/1372a85-deprecatedunsloth-Llama-3.2-3B-Instruct-model/c22b3d3-deprecatedunsloth-Llama-4-Maverick-17B-128E-Instruct-model/4d0b9b8-deprecatedunsloth-Llama-4-Scout-17B-16E-Instruct-model/afd8e49-deprecatedunsloth-Llama-4-Scout-17B-16E-Instruct-model/8de9338-deprecatedunsloth-Meta-Llama-3.1-405B-Instruct-bnb-4bit-model/4efb87c-deprecatedunsloth-Meta-Llama-3.1-70B-Instruct-bnb-4bit-model/d401984-deprecatedunsloth-Meta-Llama-3.1-8B-Instruct-bnb-4bit-model/e78288d-deprecatedunsloth-Meta-Llama-3.1-8B-Instruct-bnb-4bit-model/3da6bbe-deprecatedunsloth-Meta-Llama-3.1-8B-Instruct-bnb-4bit-model/5b0dd30-deprecatedunsloth-Phi-3-medium-4k-instruct-bnb-4bit-model/6f7b2d9-deprecatedunsloth-Qwen3.6-27B-NVFP4-model/6db1783-deprecatedunsloth-SmolLM2-135M-Instruct-GGUF-model/*-deprecatedunsloth-codegemma-7b-model/93a0a77-deprecatedunsloth-gemma-7b-it-bnb-4bit-model/5d05931-deprecatedunsloth-llama-3-70b-Instruct-bnb-4bit-model/df8fab1-deprecatedunsloth-llama-3-8b-Instruct-bnb-4bit-model/e114164-deprecatedunsloth-mistral-7b-instruct-v0.3-bnb-4bit-model/12a648d-deprecatedzai-org-GLM-4.5-Air-model/a24ceef-deprecatedzai-org-GLM-4.6-model/be72194-deprecated