Accelerated Reinforcement Learning Policy Development for Industrial Robot Manipulation Using NVIDIA Isaac Lab
This project investigates model-free and policy-optimization reinforcement learning within NVIDIA Isaac Lab to accelerate skill acquisition for robotic manipulators. Students will optimize sample efficiency and convergence rates through advanced reward shaping, discrete/continuous action-space formulations, and curriculum learning paradigms, ultimately establishing a robust, repeatable sim-to-real workflow for adaptive industrial manufacturing systems.
Required Knowledge & Skills: Familiarity with reinforcement learning fundamentals, including policies, rewards, observations, actions, and exploration; proficiency in Python and PyTorch; experience with NVIDIA Isaac Sim, NVIDIA Isaac Lab, or similar robotic simulation environments; understanding of robotic manipulators, coordinate systems, kinematics, and gripper control; experience with reinforcement learning frameworks such as RSL-RL, RL-Games, or SKRL is an asset; familiarity with reward shaping, curriculum learning, domain randomization, and sim-to-real concepts is beneficial.
Application Instructions: Interested applicants should combine their CV, undergraduate transcripts, and graduate transcripts into a single PDF file and email it directly to the supervisor. This project is conducted in collaboration with an industry partner, offering strong potential for future co-op opportunities.
Supervisor: Professor Ladan Tahvildari
3D Vision for robot planning
As a robot manipulates 3D objects and navigates 3D scenes, it requires spatial reasoning to ensure safe planning. Recent advances in 3D scene representation, such as Gaussian Splatting, enable the creation of high-fidelity digital twins of real-world environments from multi-view images. This project leverages 3D vision-language fields for open-vocabulary robot planning. (Knowledge of robotics, machine learning, and tools such as PyTorch is required.)
Supervisor: Roya Firoozi
Email: roya.firoozi@uwaterloo.ca
Deployment and testing of the SlothBot W for long-term environmental monitoring
This project aims to stress-test the SlothBot W for long-term monitoring applications. The SlothBot W is a solar-powered wire-traversing robot equipped with environmental sensors and a recharging station for aerial robots. The functionalities to be tested are active solar power harvesting, energy transfer, as well as environmental data collection, storage, and processing.
Requirements:
- Prior experience with coding in C++, Python, and ROS2
- Prior experience with single board computers and microcontroller programming
- Prior experience with CAD of mechanical and electronic components
- [Possibly] Prior experience with front-end web development
Supervisor: Prof. Gennaro Notomista
Email: gennaro.notomista@uwaterloo.ca
Location: E5 / PSE / outdoor
Fill-in-the-blanks for traffic scenarios via conditional generative models
Description: Develop and implement generative models to synthesize traffic scenarios given a partial scene (e.g., known trajectories for only 2 out of 10 vehicles in a given road map) by reasoning over unobserved behaviors to “fill in” the scenario.
Expected outcome: A robust framework that generates realistic, socially compliant, and kinematically feasible traffic scenarios given a map and partial scene (traffic scenario); re-useable codebase.
Milestones:
- Modify an existing traffic scenario model, PathDiffuser[1], to generate a traffic scenario conditioned on known partial information.
- Develop a framework to assign likelihoods to “filled-in” scenarios based on the training dataset.
Reference:
[1] D. S. Lee, A. Karthikeyan, Y. V. Pant and S. Fischmeister, "Path Diffuser: Diffusion Model for Data-Driven Traffic Simulator," 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC), Gold Coast, Australia, 2025, pp. 569-576, doi: 10.1109/ITSC60802.2025.11423013.
https://ieeexplore.ieee.org/document/11423013
Requirements:
- Knowledge and hands-on experience with generative models such as diffusion or flow.
- Proficiency in Python, Pytorch, and Slurm
Contact: Prof. Yash Pant
Geometric nonlinear control of underactuated mechanical systems
We will design and implement a nonlinear feedback controller for motion control of a rotational inverted pendulum. We will use the tools of nonlinear control and differential geometry to motivate our design and mathematically prove its effectiveness. The task includes (1) modelling the system (2) analyzing the resulting model (3) design and simulate a path following controller to move the pendulum in a desired manner (4) implement the controller on a Quanser designed hardware platform.m.
Supervisor: Prof. Chris Nielsen
Email: cnielsen@uwaterloo.ca
LLM-based multi-robot communication
Communication is both a crucial and a highly non-standardized component in coordinated multi-robot control. This project will explore the use of LLMs for natural language processing in order to realize a novel, robust, and easy-to-set-up communication infrastructure for multi-robot systems. The framework will be tested on a team of ground mobile robots in the RoboHub to perform a variety of coordinated control tasks.
Requirements:
- Prior experience with coding in Python and ROS2
- Prior experience with coding using OpenAI’s, Google’s, Meta’s or similar LLMs
- Prior experience with control of mobile robots
Supervisor: Prof. Gennaro Notomista
Email: gennaro.notomista@uwaterloo.ca
Location: E5 / PSE
LLM-Assisted Production Intelligence and Closed-Loop Control via Digital Twin Environments
This project develops an LLM-based intelligent assistant integrated with a virtual digital twin to streamline industrial automation control. By combining synthetic operational data streams with technical documentation and SOPs, the system analyzes trends, diagnoses anomalies, and executes closed-loop parameter adjustments to significantly reduce operator cognitive load and optimize overall manufacturing performance.
Required knowledge & skills: Familiarity with large language models and prompt engineering; experience with retrieval-augmented generation (RAG) and document-grounded QA systems; understanding of automation systems, PLC/SCADA concepts, and production workflows; proficiency in Python; experience with data analysis and time-series production data; knowledge of API integration and orchestration frameworks (e.g., LangChain, LlamaIndex, or similar); familiarity with digital twin platforms is an asset.
Application Instructions: Interested applicants should combine their CV, undergraduate transcripts, and graduate transcripts into a single PDF file and email it directly to the supervisor. This project is conducted in collaboration with an industry partner, offering strong potential for future co-op opportunities.
Supervisor: Professor Ladan Tahvildari
Real-Time Control and Embedded Systems for Motion Fidelity Driving Simulators
This project aims to develop a high-fidelity driving simulator that prioritizes motion cues to maximize driver retention and recall in training scenarios. A six-degree-of-freedom motion platform with linear actuators will reproduce translational and rotational dynamics, driven by real-time vehicle physics. Embedded controllers will regulate actuator signals, ensuring accurate and synchronized motion feedback. The simulator scenarios will focus on developing the telemetry within hazardous, and intentionally high-risk driving conditions that cannot be safely practiced in traditional training, such as braking on ice, hydroplaning, and collision avoidance. The project will also design a data acquisition pipeline for logging vehicle dynamics, actuator states, and driver biometrics emphasizing communication interfaces, signal processing, and embedded control systems.
Supervisor: Prof. Oliver Schneider
Email: oliver.schneider@uwaterloo.ca
Phone: 519-888-4567 x38505
Location: CPH 3627
Supervisor: Prof. Siby Samuel
Email: siby.samuel@uwaterloo.ca
Phone: 519-888-4567 x37656
Location: EC4 2119
Robot learning to adapt to user preferences
Robots are being deployed in human-centric environments, working alongside and with humans. However, different people have different preferences on how a robot should act -- how fast it should move, how close it can come, and how it should interact. These preferences are user specific, and so the robot should learn them online, and adapt in real-time. This project will build on our recent work to develop learning algorithms for robots to adapt to human preferences and improve robot performance.
Supervisor: Prof. Stephen L. Smith
Email: stephen.smith@uwaterloo.ca
Project:
Self-driving vehicles are becoming reality. Thanks to technological advances in communication and in artificial intelligence, self-driving vehicles are becoming a much closer reality than we think. Indeed, recognizing the undisputed promise of self-driving to improve safety and road efficiency, several cities around the globe have started to conduct self-driving trials in order to ensure their readiness for deploying this revolutionary technology. Nonetheless, self-driving vehicles currently face several challenges: They are limited in their speed performance, crash prevention, sensing, cooperation, and coordination capabilities. Their decision making, and hence their actions, rely merely on their on-board sensory and internal control models. To avoid crashes and prevent traffic congestion, self-driving vehicles must anticipate the behavior of other vehicles in their environment (self-driving and human-operated vehicles), share their internal state with these vehicles, and cooperatively operate to choose joint safe and efficient control policies. A primitive example of self-driving cooperative behavior can be seen in platooning, where a small group of trucks assemble in a linear structure, using wireless connectivity to maintain a prescribed distance between each other. However, more research is needed to investigate and develop techniques that enable full autonomy performance in irregular vehicle arrangements such as convoys, and in unstructured driving situations. Moving from structured small-scale working environments to large-scale unstructured working environments presents a major challenge to the future of self- driving technology due to the lack of proper understanding and modeling of the interaction dynamics in such situations; as well as due to the lack of collective sensing and cooperating methodologies that can facilitate cooperation in complex driving situations. Such collaboration is potentially feasible, given the emerging communication technology (e.g., IEEE 802.11p standard) and V2X communications services under the future 5G wireless standard. In this context, the goal of this research program is the development of a framework for coordinating interaction and cooperation among a convoy of self-driving vehicles, operating in complex driving conditions in the presence of human-operated vehicles.
1. Behaviour Modeling of Human-Driven Vehicles
Develop a class of models that can mimic the behavior of human-operated vehicles under various driving conditions- environmental, infrastructure, and traffic, and for mitigating the impact of human-operated vehicles on self-driving performance, in general, and safety in particular.
2. Situation Assessment in Human/Machine Driven Environments
Develop strategies for situational assessment, both at the vehicle level and the convoy level.
3. Self-Driving Cooperation Strategies
Develop tactical self-driving strategies that can enable interaction and cooperation between self-driving vehicles to facilitate joint mobility control and coordination, in the presence of human-operated vehicles.
4. Cooperative Resource Sharing in Self-Driving Applications
Develop self-driving strategies that can enable cooperation between convoys (coalitions) of self-driving vehicles to facilitate optimal trip planning and infrastructure resources sharing.
Note: Various group/collective behavior use-case studies should be conducted to validate the outcome of each project, both theoretically and using simulations. Examples of use cases of interest include speed harmonization, emergency response, intersection negotiation, crash avoidance, platoons and Convoys.
Robots painting music: User studies with painters and musicians
“Dark sounds” and “Dissonant colors” are just two of the many expressions highlighting the tight relationship between music and painting. The goal of this project is to connect these forms of art and human creativity via robot motion. User studies will be designed to understand how painters and musicians react to and interact with robots painting music. The project will make use of a team of ground mobile robots in the RoboHub.
Requirements:
- Prior experience with coding in Python and ROS2
- Prior experience with control of mobile robots
- [Possibly] Prior experience with design of user studies
Supervisor: Prof. Gennaro Notomista
Email: gennaro.notomista@uwaterloo.ca
Location: E5 / PSE
Teleoperation of teams of wheeled drones
This project aims to develop shared-control algorithms for teams of wheeled drones teleoperated by a human. The activities will include the programming of optimization-based robot controllers in ROS2 and their tests using a haptic device and wheeled drones in the RoboHub.
Requirements:
- Prior experience with coding in Python and ROS2
- Prior experience with control of mobile robots
- Good knowledge (theory and implementation) of control theory and mathematical optimization
- [Possibly] Prior experience with CAD of mechanical components
Supervisor: Prof. Gennaro Notomista
Email: gennaro.notomista@uwaterloo.ca
Location: E5 / PSE