Job Description
About the Role
The AI for Science Gym builds bottom-up AI capability across NUS science and engineering. Underneath its training gyms sits the Instrumentation Gym — the shared substrate of curated datasets, trained models, reusable workflows, and the harness that connects them on NUS HPC.
You will design and build that substrate's foundational layers. These are the entry points through which most researchers engage with the Gym, and the patterns you set determine how cleanly the system scales onto multi-node HPC. The role sits at the interface between scientific data, ML representations, and the compute environments researchers use day-to-day.
Job Description
- Data pipelines — ingestion, metadata schemas, and provenance tracking for electron microscopy, X-ray, and light microscopy datasets; format normalisation so one lab's deposit is usable by another.
- Model interfaces — modular Python APIs that let representation models, tokenizers, and downstream models be swapped and composed without rewriting the pipeline around them.
- Baseline tokenization and cartography pipelines, built collaboratively with ML researchers and domain scientists, as reference implementations others adapt for new domains.
- LLM- and RAG-assisted interfaces for dataset navigation, workflow discovery, and user interaction — retrieval over the Gym's own datasets, models, workflows, and deposited challenges.
- Forward-compatibility to HPC — components that scale from single-GPU teaching instances to multi-node; containerised, reproducible, scheduler-friendly deployment on NUS HPC.
- Onboarding and enablement — semi-regular sessions for new Gym users; helping instructors build workflows that generate teaching material from their own sources plus Gym content.
Qualifications
Degree in a Computational, Engineering, or Physical-Science discipline. A Master's or Bachelor's with equivalent research-engineering experience.
Skills:
- Strong Python and a modern ML framework (PyTorch preferred).
- Experience building data pipelines and reproducible ML workflows end to end, not just training scripts.
- Comfort with GPU compute and scientific data formats (HDF5, TIFF/OME-TIFF, Zarr, or instrument-native equivalents).
- Track record of modular, maintainable software — tested, documented, and successfully picked up by people who didn't write it.
- Interest in working at the intersection of science, ML, and systems engineering, and the temperament to stay there.
- Ability to communicate with non-engineers — much of the job is understanding what a scientist actually needs.
- Scientific imaging data — electron microscopy, X-ray/synchrotron, or light microscopy, including its artefacts and metadata conventions.
- LLMs and retrieval-augmented generation: embeddings, vector stores, and the failure modes of retrieval over technical corpora.
- Interactive data tools — notebook UIs, dashboards, viewer or annotation apps.
- Deploying ML services on HPC or cloud GPU infrastructure: Slurm, Docker/Apptainer, distributed training.
- Representation learning, self-supervised or foundation models, or tokenization for non-text modalities.
- Open-source contributions in scientific Python or ML tooling; FAIR data practice.