~/projects/ai-simulation-platform
Reinforcement Learning Simulations
Learning through observations and rewards
UnityC#PythonML-AgentsPPOPyTorch
$ overview
I wanted to understand how much an agent's behavior depends on what it can observe and what it gets rewarded for. I started with food-seeking agents without vision, then added static and moving obstacles, followed by visual observations.
I built the environments in Unity ML-Agents, with modular C# logic and Python pipelines for PPO training. Each new task gave me a reason to revisit the observations, reward functions, and training settings. I later extended the experiments to a six-axis robotic arm, which introduced a more complex action space and control problem.
$ implementation notes
- Progressed from food-seeking without vision to navigation around static and moving obstacles, then experiments with visual observations.
- Reworked observations and reward functions through successive training runs, using C# environment logic and Python training pipelines.
- Extended the simulations to a six-axis robotic arm to explore a more complex set of observations, actions, and rewards.