Robin Nicole, PhD · ML/AI Engineer · London

Ship AI to production.

Senior ML engineer. RAG, agents, forecasting, MLOps. Production systems that run for years.

See the work

No hand-holding needed. Brief in, working system out. I speak the language.

RAG concierge with retrieval grounding, cost guards, and an eval harness. Doubled ops throughput.

Selected work

Three systems, shipped and still running

RAG concierge

Problem
Customer operations was the bottleneck. Every incoming question required manual search across scattered docs, past tickets, and dashboards. Latency high, throughput flat, no path to scale with hiring.
Build
RAG concierge end-to-end: document ingestion pipeline, vector retrieval over internal KB, retrieval-grounded LLM responses with inline citations, cost and safety guardrails, and an eval harness tracking answer faithfulness, citation coverage, and drift as content changed.
Outcome
2x operational throughput without headcount. Query latency dropped from minutes of manual search to sub-second retrieval. Still running in production, on-budget, within SLO.

Demand forecasting at retail scale

Problem
Manual and coarse demand forecasts across European retail markets meant overstock in some places and stockouts in others, expensive both ways, with no way to scale the approach across markets and SKUs by hand.
Build
Demand forecasting and inventory pipelines across European markets on VertexAI and Kubeflow, BigQuery as the data layer. Hierarchical time-series models, automated retraining, and drift monitoring so forecasts stay honest as demand shifts.
Outcome
Forecasts running in production across multiple European markets, retrained automatically and monitored for drift, replacing manual estimates with a system the team can trust without babysitting.

Real-time NLP at a national publisher

Problem
One of the UK's largest news publishers needed to classify and route editorial content at scale, in real time. Manual tagging was too slow for the volume and editors had no reliable way to test what worked.
Build
NLP content-classification system over real-time data pipelines handling 100M+ monthly page views, plus an A/B testing framework built to be simple enough that editors actually used it day to day.
Outcome
Classification running in real time across 100M+ monthly page views; the A/B framework became part of the editorial workflow rather than a dashboard nobody opened.

FAQ

Common questions

Q & A
What's your stack?
Python first. TensorFlow, PyTorch, scikit-learn, pandas, PySpark. GCP (VertexAI, Kubeflow, BigQuery) and AWS (EMR, S3). Docker. For LLM work: bare SDKs or LangChain/LlamaIndex depending on constraints, vector DBs (pgvector, Qdrant, Pinecone) as needed, eval harnesses always. I'm stack-flexible inside this range.
Q & A
Do you handle MLOps or just modelling?
Both, and in that order. I care as much about monitoring, CI/CD, feature stores, eval infra, and incident retros as I do about the model. Modelling without operational discipline is how companies end up with models that never ship.
Q & A
How do you evaluate LLM outputs?
Layered. Offline eval sets (golden answers, LLM-as-judge where appropriate), online metrics with human spot checks, drift detection on content changes, cost/latency SLOs. For RAG specifically: retrieval recall on labelled queries, answer faithfulness, citation coverage. Evals ship with the system, not after.
Q & A
How does an engagement usually start?
Short intro call to scope the problem. Proposal with clear deliverables and timeline. Embedded with your team or working independently depending on what fits. Weekly review calls, handoff docs written for the engineer who takes over, not for show.

Services

What I ship

Service

LLM & RAG systems, built to ship

Retrieval-augmented chatbots, agents, and production LLM features. Prompt engineering, eval harnesses, cost controls, and the MLOps to keep them running. Python with SDKs or LangChain/LlamaIndex depending on constraints, vector DBs as needed (pgvector, Qdrant, Pinecone), cloud or self-hosted.

2 weeks to 3 months

Service

Forecasting & demand modelling

Demand forecasting and optimization pipelines at scale. Time-series models and hierarchical forecasting. Productionised on VertexAI/Kubeflow or your platform of choice.

4 weeks to multi-month embedded

Service

ML platform & MLOps

Production-grade pipelines, monitoring, CI/CD for models, feature stores, eval infra, incident hygiene. Systems that land in prod and stay there.

Retainer or scoped platform build

Experience

Prior work

Track record

Seven years across retail, adtech, media, and quantitative finance. Production ML that runs for years.

  1. Oct 2024, Present

    Senior ML Engineer · Major European retailer

    Demand forecasting across European retail markets. VertexAI and Kubeflow pipelines, BigQuery, Python.

  2. Jan 2022, Sept 2024

    Senior Applied Scientist · Sojern

    ML for travel advertising. LLM-powered concierge (RAG chatbot). ML pipeline automation, eval harnesses, cost controls on prod LLM workloads.

  3. Sept 2019, Jan 2022

    Senior Data Scientist · Reach plc

    NLP systems for one of the UK's largest news publishers. Content classification, real-time data pipelines, A/B testing framework editors actually used.

  4. May 2018, Sept 2019

    Quantitative Analyst · Gambit Research

    ML models for market prediction. Signal research and backtesting in a quant trading environment.

About

About

Senior ML engineer. PhD Applied Mathematics, King's College London. Seven years of production ML across retail demand forecasting, adtech bidding, news NLP, and quant prediction.

Most of what I ship is systems that run for years, not demos. Demand forecasting and inventory across European retail. RAG concierge doubling ops throughput for a client. NLP classification over 100M+ monthly page views at a national publisher. Quant signals at a market-making firm.

I care about systems that survive. That means monitoring, eval harnesses, honest incident retros, and modelling choices that let the next engineer on-call sleep. I'll move fast, but I won't ship something I wouldn't want to own.

Publications

Stack

ML/AI
Deep Learning, NLP, LLMs, Agentic AI, Time Series, RAG
Languages
Python, SQL, C++
Frameworks
TensorFlow, PyTorch, scikit-learn, pandas, PySpark
Cloud/MLOps
VertexAI, Kubeflow, Docker, AWS, GCP

Contact

Get in touch

Brief in, reply within a week. If I'm not the right fit, I'll usually know who is.

Include: the problem, the data shape, your existing stack, and a rough timeline. No preamble needed.