About
I am a Master's student at the University of Science and Technology of China (USTC),
focusing on intelligent systems that combine retrieval, reasoning, and adaptive memory.
My current research centers on personalized agentic RAG systems with autonomous
decision-making and self-correction loops.
I am also deeply interested in reinforcement learning for LLM agents,
long-tailed recognition, and sharpness-aware minimization (SAM)
methods for imbalanced datasets. I enjoy building end-to-end systems and open-sourcing
my work when possible.
Research Vision & PhD Goals
What I Want to Build
I aim to develop embodied agents that learn continuously from interaction β
systems that combine the reasoning power of large language models with the grounded
decision-making of reinforcement learning. My vision is agents that can adapt to new
environments, correct their own mistakes, and transfer knowledge across tasks without
catastrophic forgetting.
My Master's work on agentic RAG and multi-agent orchestration has convinced me that the next
frontier is not just better models, but better agents β systems that can
perceive, reason, act, and learn in open-ended environments. A PhD will give me the depth
and mentorship to pursue this at the highest level.
Specific Research Directions
-
RL for LLM Reasoning: Train language models to reason through
verifiable rewards (GRPO, PPO) rather than supervised fine-tuning alone β enabling
self-improving agents for math, code, and planning.
-
Continual & Lifelong RL: Design agents that learn new tasks sequentially
without forgetting prior knowledge, using techniques like EWC, progressive networks, and
episodic memory.
-
Embodied AI & Robotics: Bridge high-level LLM reasoning with low-level
robotic control through Vision-Language-Action (VLA) models and skill discovery.
-
Multi-Agent Coordination: Extend my Agent Prime work to decentralized
multi-agent systems with economic incentives, reputation, and emergent collaboration.
I am actively looking for PhD advisors working in these areas, particularly at institutions
with strong RL and robotics groups. I am open to positions in mainland China, Hong Kong,
and internationally.
Research Interests
Agentic AI Systems
Retrieval-Augmented Generation (RAG)
Reinforcement Learning
Long-Tailed Learning
Sharpness-Aware Minimization
Multimodal Learning
LLM Agents
Embodied AI
Selected Projects
-
Personalized Agentic RAG System
2025 - Present
Designing memory-augmented agentic RAG with adaptive retrieval, user modeling,
confidence estimation, and self-correction loops. Built with LangGraph, LangChain, FAISS, and LLMs.
-
Campus Knowledge Base RAG Chatbot
2025
A RAG prototype for campus knowledge retrieval using Ollama embeddings,
Chroma vector database, and a Gradio-based user interface.
-
Class-Adaptive Focal SAM
Master's Thesis, 2025β2026
A class-adaptive focal loss combined with Sharpness-Aware Minimization (SAM)
for long-tailed visual recognition. Evaluated on CIFAR-10/100-LT, TinyImageNet,
ImageNet-LT, and iNaturalist 2018.
Thesis β Code coming soon
-
Agent Prime - Multi-Agent Orchestration with USDC Nanopayments
Lablab.ai Agentic Economy on Arc Hackathon, 2026
A 3-agent workflow (Research / Writer / Editor) with real USDC nanopayments
on Arc Testnet. Includes on-chain verification via web3.py and hits both
Usage-Based Compute Billing and Agent-to-Agent Payment Loop tracks.
-
Resume / Candidate Screening System
Future Interns - Remote Internship, MarβApr 2026
An ML-powered resume screening and candidate ranking system built during
a remote internship. Uses NLP and transformer-based models for automated candidate evaluation.
-
RL Portfolio - Dueling DQN, PPO, A2C, VPG
Reinforcement Learning / 2026
Hands-on implementation of Dueling DQN and modern RL algorithms
based on ARENA, UC Berkeley CS 285, and Spinning Up.
Exploring RL for LLM agents and mechanistic interpretability.
-
Weather Trend Forecasting
Time Series / ML / 2025
Weather trend forecasting using time series analysis and ML models.
-
AI Study Coach
EdTech / NLP / 2024
An interactive AI-powered study assistant with personalized learning paths,
quiz generation, and progress tracking using LLM-based techniques.
Blog
-
From Tutorial to Production: Building a Modular RLHF Pipeline
July 2026
A deep dive into transforming three tutorial scripts (SFT β Reward Model β PPO) into a
production-grade, configurable RLHF framework. Covers Bradley-Terry loss, clipped PPO objectives,
KL divergence penalties, and lessons learned from training a 0.5B model end-to-end.
Publications
Coming soon β publications in preparation.
Experience
-
Mar β Apr 2026
Machine Learning Intern
Future Interns (Remote)
Built a Resume/Candidate Screening System using NLP and transformer models.
-
Nov 2022 β Aug 2023
IT Assistant
AAMUSTED - Akenten Appiah-Menka University of Skills Training and Entrepreneurial Development
Provided IT support, system maintenance, and technical assistance.
Education
-
2024 β 2027 (Expected)
M.S. in Software Engineering (AI/ML)
University of Science and Technology of China (USTC)
-
2018 β 2022
B.Sc. in Information Technology Education
University of Education, Winneba
First Class Honors
Technical Skills
Languages: Python, JavaScript, SQL, Java, Bash
Frameworks & Libraries: PyTorch, scikit-learn, Transformers (Hugging Face), LangChain, LangGraph, AutoGen, FAISS, Chroma, spaCy, NLTK
Tools & Platforms: Git, Docker, Linux, wandb, FastAPI, Streamlit, Jupyter, VS Code, LaTeX
Cloud & GPU: AutoDL, Google Cloud Shell, NVIDIA AI Endpoints, Remote GPU servers (4090/5090-class)
Certifications
-
NVIDIA - LLM RAG Agent Fundamentals (2025)
-
DeepLearning.AI - Advanced Learning Algorithms (2025)
-
DeepLearning.AI - Supervised Machine Learning (2025)
News
Jul 2026
Applied to HKUST CSE Early Recruiting for Fall 2027 PhD intake. Awaiting interview decision.
May 2026
Participated in the Lablab.ai Agentic Economy on Arc Hackathon - built Agent Prime with real USDC nanopayments.
Apr 2026
Completed remote ML internship at Future Interns with a Resume/Candidate Screening System project.
2025
Started Master's research on Class-Adaptive Focal SAM for long-tailed visual recognition.