01 / ABOUT

I'm a Master's student in Computer Science at Northeastern, focused on GenAI — RAG pipelines, agents, and the eval harnesses and guardrails that decide whether an LLM feature is trustworthy enough to ship.

Underneath that is a backend engineer: I build the distributed systems — Go services, AWS infra, the production plumbing — that turn a model into something people can actually rely on.

When I'm not at a keyboard, I'm probably on a chairlift at Whistler, behind a camera, or going way too deep into NASA space documentaries.

02 / EDUCATION
2024 — Dec 2026

MS, Computer Science · Northeastern University

Focused on machine learning and the cloud infrastructure that AI systems run on. Coursework: Machine Learning, Scalable Distributed Systems, Database Management Systems, Algorithms.

2020 — 2024

B.Econ, Finance · Minzu University of China

Pivoted into software toward the end of undergrad; that finance background still shapes how I think about engineering tradeoffs.

View Full Résumé ↗
03 / PROJECTS
06 shown
[ Bedrock RAG · 2-region AWS ]
AIBACKEND

CanPlan — RAG Task Planner

A RAG backend on AWS that turns a daily-living goal into grounded, step-by-step plans, each citing its source. Three measured eval rounds on CanPlan's step-generation service ended at 100% valid citations and a reading level cut from grade 7.2 to 6.0 — and showed the safety guardrail still over-flags (96% → 89% of plans).

TypeScriptAWS CDKBedrockS3 VectorsDynamoDBRAG Eval
[ F1 0.83 · 1.31% params ]
AIMACHINE LEARNING

Financial Sentiment Analysis & Agent

↗

Fine-tuned a small language model to read financial text as bullish, bearish, or neutral — F1 0.83 while training only 1.31% of params (a 3.4 MB adapter) — then wrapped it in a chat agent that pulls live headlines. On live news it scores 0.66, level with FinBERT, and a fixed pipeline matched the agent's reliability with 7× fewer tokens.

Live Demo↗
PythonLLMFine-tuningLangChain
[ ~1,884 req/s · p95 −21% ]
BACKEND

Distributed Order Processing System

↗

An e-commerce order backend split into small services that never oversell stock, even under load — 599 concurrent orders on 200 units settle to exactly 200. It sustains ~1,884 req/s at 56 ms p95, and a connection-pool tuning round cut p95 21% (140 → 110 ms).

GoGinRabbitMQAWSPrometheusGrafanaDocker
[ 0.185 MB · F1 0.943 ]
MACHINE LEARNING

VAD Model Compression

↗

Mapped the size/F1 Pareto front for a CRDNN speech-detection model from 0.435 MB down to 0.050 MB. The shippable point is 0.185 MB at F1 0.943, against a 0.959 baseline — squeezing to 0.050 MB is 8.7× smaller but costs 9 points of F1. Dynamic quantization barely moved it, because only 1.2% of the parameters were quantizable.

PyTorchSpeechBrainModel Compression
[ 30GB+ ETL · 26 workers ]
BACKEND

Distributed Book Recommendation

↗

A book recommender that chews through 30 GB+ of public library data into a 50,000-book catalog, then fans queries out across 26 workers / 27 partitions — figuring out which steps actually benefit from going parallel.

PythonAWSDynamoDBECS
[ Full-stack · React · Express ]
BACKEND

Movify — Movie Social Network

↗

A full-stack movie social network where people register, follow each other, and post reviews — then like and vote on both movies and reviews. A React + Redux front end talks to an Express / MongoDB API that handles sessions and the social graph, and proxies live movie data from the TMDB API.

ReactRedux ToolkitTypeScriptExpressMongoDBTMDB API
04 / CONTACT

Let's build something great

Looking for an AI engineer starting Dec 2026 — or just want to talk RAG, agents, and ski lines? My inbox is open.

Email me →
© 2026 Jiaming Pei · Vancouver, BC