back to projects

~/projects/ai-simulation-platform

Reinforcement Learning Simulations

Learning through observations and rewards

Jan 2025 – Present·active
UnityC#PythonML-AgentsPPOPyTorch

$ overview

I wanted to understand how much an agent's behavior depends on what it can observe and what it gets rewarded for. I started with food-seeking agents without vision, then added static and moving obstacles, followed by visual observations.

I built the environments in Unity ML-Agents, with modular C# logic and Python pipelines for PPO training. Each new task gave me a reason to revisit the observations, reward functions, and training settings. I later extended the experiments to a six-axis robotic arm, which introduced a more complex action space and control problem.

$ implementation notes

  • Progressed from food-seeking without vision to navigation around static and moving obstacles, then experiments with visual observations.
  • Reworked observations and reward functions through successive training runs, using C# environment logic and Python training pipelines.
  • Extended the simulations to a six-axis robotic arm to explore a more complex set of observations, actions, and rewards.