Job Description

Job Title:  Research Analyst (AI Data Engineer)
University-Level Unit:  Asian Institute of Digital Finance
Faculty/Department-Level Unit:  Credit Research Initiative
Employee Category:  Research Staff
Location_ONB:  Kent Ridge Campus
Posting Start Date:  21/08/2026

Background

The Asian Institute of Digital Finance (AIDF) is a university-level institute in NUS, jointly founded by the Monetary Authority of Singapore (MAS), the National Research Foundation (NRF) and NUS. AIDF aspires to be a thought leader, a Fintech knowledge hub, and an experimental site for developing digital financial technologies as well as for nurturing current and future Fintech researchers and practitioners in Asia. The Credit Research Initiative (CRI) is a non-profit undertaking under the AIDF. Pioneering the “public good” credit risk measures, the initiative is committed to advancing big data analytics and providing directly useful credit intelligence to academic and professional communities.

Reliable data infrastructure sits at the core of our operations — from daily credit risk production to AI and LLM research. AIDF-CRI is dedicated to staying updated with the latest trends and technologies, and we are continually enhancing our data pipelines and applying AI to our research and workflows. We are looking for an AI Data Engineer who can take genuine ownership of this foundation and keep it evolving with the demands of the AI era.

Responsibilities

Data underpins nearly everything we do: our credit risk measures, analytics, and AI models all depend on the quality of the data behind them. Efficient and reliable data collection, preparation, and delivery are easy to overlook, yet they are critical to high-performing models and trustworthy research. At its core, therefore, this is a data engineering role in the AI era. The selected candidate will take ownership of our existing data pipelines, databases, scheduling and monitoring, and operate them reliably. Beyond this core scope, the role also offers opportunities to gain additional exposure to research-oriented work targeting top AI conferences and/or R&D work under our industrial research collaborations.

Particularly, the responsibilities will include:

  • Data Pipeline Ownership & Maintenance
    • Operate, maintain, and enhance our existing ETL/ELT data pipelines.
    • Own scheduling and workflow orchestration, including monitoring, alerting, data quality checks, and incident recovery, to keep pipelines reliable.
    • Develop and optimize data models and schemas to support analytics, reporting, and machine learning requirements.
  • Database & Infrastructure Management
    • Manage and tune our relational, NoSQL, and vector databases for storage, querying, and retrieval workloads.
    • Improve the robustness, efficiency, and cost-effectiveness of the data infrastructure as data volumes and use cases grow.
  • AI Research & Automation
    • Where involved, contribute to research projects targeting top AI conferences, and/or R&D work under our industrial research collaborations, building on deep familiarity with our data assets.
    • Occasional ad-hoc tasks applying LLM-based agents to automate parts of our data collection, document processing, and other internal workflows.
  • Team Collaboration & Documentation
    • Collaborate with financial analysts and the R&D team to ensure data accessibility and usability.
    • Maintain comprehensive documentation of pipelines, system architecture, and database schemas to promote knowledge sharing and smooth onboarding.

Minimum Requirements

  • Bachelor’s or Master’s degree in Computer Science, Information Technology, Data Science, Computer Engineering, or a related field.
  • Strong proficiency in Python and SQL, with solid software practices (e.g., Git, virtual environments, testing mindset).
  • Hands-on experience building or maintaining ETL/ELT data pipelines on cloud platforms; working knowledge of Google Cloud Platform and/or Snowflake is strongly preferred.
  • Experience with workflow orchestration and scheduling tools (e.g., Airflow / Cloud Composer), including monitoring, alerting, and troubleshooting of production data pipelines.
  • Experience working with relational and NoSQL databases (e.g., MySQL, MongoDB) — designing schemas, writing performant queries, and managing day-to-day operations.
  • Strong problem-solving skills and the ability to work independently in a fast-paced environment.
  • Excellent communication and documentation skills, both written and verbal, to collaborate effectively with cross-functional teams and stakeholders.

Bonus Skills

  • Hands-on experience integrating LLMs into data pipelines or internal workflows — e.g., RAG / vector search, fine-tuning, or agentic automation.
  • Exposure to vector databases or extensions (e.g., Milvus, PostgreSQL with pgvector).
  • Experience applying NLP to alternative data at scale, such as news articles, financial filings, and social media content.
  • Publications, open-source contributions, or a strong project portfolio; interest in pursuing research toward top AI conferences.
  • Experience with Docker, CI/CD pipelines, and cloud deployment.
  • Familiarity with machine learning techniques in a financial context
Req ID:  34171