Accelerated Reinforcement Learning Policy Development for Industrial Robot Manipulation Using NVIDIA Isaac Lab
This project investigates model-free and policy-optimization reinforcement learning within NVIDIA Isaac Lab to accelerate skill acquisition for robotic manipulators. Students will optimize sample efficiency and convergence rates through advanced reward shaping, discrete/continuous action-space formulations, and curriculum learning paradigms, ultimately establishing a robust, repeatable sim-to-real workflow for adaptive industrial manufacturing systems.
Required Knowledge & Skills: Familiarity with reinforcement learning fundamentals, including policies, rewards, observations, actions, and exploration; proficiency in Python and PyTorch; experience with NVIDIA Isaac Sim, NVIDIA Isaac Lab, or similar robotic simulation environments; understanding of robotic manipulators, coordinate systems, kinematics, and gripper control; experience with reinforcement learning frameworks such as RSL-RL, RL-Games, or SKRL is an asset; familiarity with reward shaping, curriculum learning, domain randomization, and sim-to-real concepts is beneficial.
Application Instructions: Interested applicants should combine their CV, undergraduate transcripts, and graduate transcripts into a single PDF file and email it directly to the supervisor. This project is conducted in collaboration with an industry partner, offering strong potential for future co-op opportunities.
Supervisor: Professor Ladan Tahvildari
Action Detection in Road Scenes
Leveraging the ROAD Dataset, WATonomous is developing an action classifier for road participants from video streams. The system will be deployed on-vehicle and validated in lead vehicle overtaking and pedestrian interaction scenarios. The team is looking for experience in computer vision research, development with PyTorch and Docker, and deploying neural networks to on-vehicle GPUs.
Supervisor: Prof. Derek Rayside
Email: drayside@uwaterloo.ca
An eBPF-Based Telemetry Layer for Hardware-Assisted Malware Detection
Hardware performance counters (HPCs) expose microarchitectural program behaviour—such as branch, cache, and speculative-execution activity—that is invisible to conventional software telemetry. Our group’s PRISM detector exploits these signals by scheduling complementary sets of hardware events over time. PRISM’s false-positive rate, however, remains too high for stand-alone deployment at scale. In practice, hardware telemetry must therefore serve as an early, low-cost layer within a broader behavioural-monitoring stack.
In this project, the student will build such a stack: an eBPF-based monitoring pipeline that collects system-call- and process-level behavioural signals alongside per-process HPC traces. The student will time-align the two sources of telemetry and evaluate how effectively a second contextual layer can reduce PRISM’s false-positive rate. The evaluation will use real benign and malicious workloads in a sandboxed environment provided by our group and will quantify the detection–overhead tradeoff of the combined system relative to HPC-only and eBPF-only baselines.
You will learn about: eBPF and Linux kernel instrumentation; hardware performance monitoring using perf; applied machine learning for security; and end-to-end security-system evaluation.
Required skills: Strong C/C++ and Linux systems-programming skills; proficiency in Python for data analysis and machine learning; and prior exposure to operating-system internals.
Supervisor: Prof. Seyed Majid Zahedi
Email: smzahedi@uwaterloo.ca
Causal Inference for Simulation-Based Decision Making
This project focuses on improving causal reasoning in AI systems. This could include improving causal inference via learning from simulated or real-world data. Alternatively, it could involve looking at failures of existing generative models for text or images in causal reasoning
Required Knowledge & Skills:
Prior experience with machine learning experimental methods, and implementing some deep learning models in Python-based frameworks, is expected.
Supervisor: Prof. Mark Crowley
Email: mark.crowley@uwaterloo.ca
Consistency Learning in Forest LiDAR Data
This project investigates the development of novel machine learning architectures tailored to structured spatial data, with a focus on LiDAR scans of forest environments. The student will explore how prior domain knowledge and structural inductive biases can improve learning performance, robustness, and generalization in models trained on complex 3D sensor data.
Required Knowledge & Skills:
Prior experience with machine learning experimental methods, and implementing some deep learning models in Python-based frameworks, is expected.
Supervisor: Prof. Mark Crowley
Email: mark.crowley@uwaterloo.ca
Development of Diverse Model Generation, Selection and Ensemble Tool (University-Industry collaborative project)
In data science and machine learning, leveraging a single model often limits predictive accuracy and robustness, especially in complex, high-dimensional datasets. Model diversity through ensemble techniques—such as bagging, boosting, and stacking—can significantly enhance performance by combining the strengths of multiple models. This project focuses on developing a Diverse Model Generation, Selection, and Ensemble Modeling Tool to automate and optimize the process of model creation, evaluation, and ensembling for a wide range of applications.
The tool will generate a pool of diverse machine learning models, evaluate their performance on specified metrics, and select the best-performing models. Through ensemble techniques, it will then combine these models to achieve improved accuracy, stability, and generalization across datasets. Built with flexibility and scalability in mind, the tool will be suitable for various types of machine learning tasks, from regression and classification to time series forecasting and beyond.
Must have skills/courses by a candidate to conduct the project:
Experience with Python programming is required. Candidates with experience in machine learning (e.g. Random Forest, Gradient Boosting, eXtreme Gradient Boosting, LightGBM, CatBoost) would be given higher preference.
Interested students may send the following documents in the form of a single PDF file to the two supervisors at snaik@uwaterloo.ca and marzia.zaman@uwaterloo.ca:
- CV
- Undergraduate Transcript
- Graduate Transcript
Supervisor: Prof. Kshirasagar Naik
Email: snaik@uwaterloo.ca
Co-Supervisor: Dr. Marzia Zaman
Email: marzia.zaman@uwaterloo.ca
Development of Machine Learning Method to Perform Multi-class Classification with Imbalanced Data for Fault Diagnosis of Aero-engine Components (University-Industry collaborative project)
In the aviation industry, accurately diagnosing faults in aero-engine components is critical for ensuring safety, reducing downtime, and minimizing maintenance costs. However, identifying faults is challenging due to the imbalanced nature of the data, where some fault types are rare but highly consequential. This project aims to develop a machine learning method specifically designed for multi-class classification with imbalanced data to reliably diagnose faults in aero-engine components.
The project will involve building a model capable of accurately classifying faults across multiple categories, addressing data imbalance by employing techniques like synthetic data generation, cost-sensitive learning, and specialized algorithms tailored for imbalance. The final deliverable will be a fault diagnosis tool that can distinguish between multiple types of faults and non-faulty states, supporting preventive maintenance and enhancing operational safety.
Must have skills/courses by a candidate to conduct the project:
Experience with Python programming is required. Candidates with experience in machine learning, deep learning (e.g. LSTM, CNN, RNN, U-Net) would be given higher preference.
Interested students may send the following documents in the form of a single PDF file to the two supervisors at snaik@uwaterloo.ca and marzia.zaman@uwaterloo.ca:
- CV
- Undergraduate Transcript
- Graduate Transcript
Supervisor: Prof. Kshirasagar Naik
Email: snaik@uwaterloo.ca
Co-Supervisor: Dr. Marzia Zaman
Email: marzia.zaman@uwaterloo.ca
Development and validation of Data-driven Biomass Prediction Models with LiDAR and Satellite images (University-Industry collaborative project)
Accurate biomass estimation is essential for monitoring forest health, assessing carbon stocks, and informing sustainable land management practices. Traditional biomass measurement methods are often labor-intensive and limited in spatial scope, making it challenging to assess large, forested areas efficiently. This project aims to develop data-driven machine learning models for biomass prediction that leverage Lidar and satellite imagery to provide accurate, large-scale biomass estimates.
The project will involve using Lidar data, which provides precise structural information about vegetation, and multispectral satellite images that capture broader spatial and spectral information, to develop models capable of estimating biomass at various scales. Machine learning algorithms will be trained and validated on these remote sensing datasets to create models that can accurately predict biomass across diverse landscapes and environmental conditions. This will result in a scalable, efficient biomass estimation tool for applications in environmental monitoring, carbon budgeting, and conservation.
Must have skills/courses by a candidate to conduct the project:
Experience with Python programming is required. Candidates with experience in data processing, machine learning, deep learning (e.g., CNN, RNN, U-Net) would be given higher preference.
Interested students may send the following documents in the form of a single PDF file to the two supervisors at snaik@uwaterloo.ca and marzia.zaman@uwaterloo.ca:
- CV
- Undergraduate Transcript
- Graduate Transcript
Supervisor: Prof. Kshirasagar Naik
Email: snaik@uwaterloo.ca
Co-Supervisor: Dr. Marzia Zaman
Email: marzia.zaman@uwaterloo.ca
Ethical Reasoning in Reinforcement Learning Agents
This project explores how reinforcement learning agents can be designed to exhibit ethical decision-making in environments that simulate moral dilemmas. The student will investigate frameworks for embedding moral principles or meta-rules into learning systems and evaluate agent behavior in simulated contexts such as Minecraft.
Required Knowledge & Skills:
Prior experience with machine learning experimental methods, Deep Reinforcement Learning, and implementing some deep learning models in Python-based frameworks, is expected.
Supervisor: Prof. Mark Crowley
Email: mark.crowley@uwaterloo.ca
Exploring new avenues for accelerating vector similarity search
In data science and machine learning, a “vector” (also called an “embedding”) represents the “learned” latent space features of an item or entity in a collection of structured or unstructured data, e.g., image, document, audio and knowledge graph. Identifying similar vectors in a collection (e.g., a vector database) is a fundamental operation with wide applications to machine learning workflows, recommender systems and data mining. As the dimension of vectors become large (e.g., 10^2 – 10^3) or the size of the collection grows (e.g., 10^9 – 10^12), vector similarity search becomes computationally challenging, even more so when high accuracy is sought. This project will explore systems techniques to accelerate vector similarity search on parallel computing platforms towards improving performance at scale. More precisely, this project will focus on the top-k nearest neighbor search problem in a vector data collection. There exist several algorithmic techniques for nearest neighbor search in vector data collections: product quantization, local sensitivity hashing, clustering-based, e.g., DBSCAN and graph-based, e.g., HNSW; they offer different trade-offs in terms of throughput, scalability and accuracy. Additionally, past studies showed that it is possible to combine multiple algorithmic techniques in the same task to achieve further improvements. This project will capitalize on the existing algorithmic techniques and aims at engineering a parallel/distributed solution that offers high throughput and accuracy.
Required knowledge & skills: Object-oriented programming, data structures, familiarity with parallel/distributed computing concepts and implementation frameworks (e.g., CUDA, OpenMP, MPI, Hadoop/Spark), knowledge of machine learning concepts and tools (e.g., PyTorch, TensorFlow).
Supervisor: Prof. Ladan Tahvildari
Email: ladan.tahvildari@uwaterloo.ca
Supervisor: Prof. Tahsin Reza
Email: tahsin.reza@uwaterloo.ca
Fair GPU Scheduling for Multi-Tenant LLM Serving
Shared GPU clusters for LLM serving present a fairness challenge that classical per-round allocation mechanisms cannot address: a tenant that underuses its share during periods of low demand receives no additional priority when its demand later surges. Our recent work introduces credit fairness, an accounting-based notion of fairness for repeated resource allocation in which the mechanism keeps track of which tenants have lent or borrowed capacity over time.
The goal of this project is to conduct a feasibility study that translates credit fairness from theory into a real serving stack running on AMD GPUs. The student will map credit-based accounting onto token-level serving metrics, such as GPU memory usage, KV-cache occupancy, and batching slots; implement a credit-aware admission and scheduling policy in vLLM’s ROCm backend on AMD Instinct GPUs—or in a faithful discrete-event simulator calibrated using ROCm measurements; and evaluate the utilization–fairness tradeoff under bursty, heterogeneous, multi-tenant workloads. The proposed policy will be compared with max-min fairness and token-level fairness baselines.
Working with the ROCm stack is a deliberate choice. Its open-source runtime and kernel driver, including the AMDGPU/KFD queue and memory-management layers, expose scheduling and accounting internals that are closed on other vendors’ platforms. This transparency is important because the project’s central question is how accurately a scheduler can measure each tenant’s actual resource consumption.
The project will culminate in a written report characterizing when credit fairness is beneficial, its effect on utilization, and the operational realities—such as sequential decoding and variations in effective capacity—that challenge the underlying theory.
You will learn about: fair division and mechanism design for computer systems; LLM inference-serving internals, including continuous batching and KV-cache management; the AMD ROCm software stack, including HIP, rocprof, and vLLM’s ROCm backend; and workload-driven performance evaluation.
Required skills: Strong Python and/or C++ skills and experience with systems programming. Familiarity with GPU programming—through HIP/ROCm or transferable CUDA experience—or with machine-learning serving systems is an asset.
Supervisor: Prof. Seyed Majid Zahedi
Email: smzahedi@uwaterloo.ca
Fair Sharing of an Open-Source FPGA GPU: Vortex-Based LLM Inference on Stratix 10
This project will build a hardware testbed for our research on fair GPU sharing, complementing Project 1, using Vortex: [https://github.com/vortexgpgpu/vortex](https://github.com/vortexgpgpu/vortex). Vortex is a full-stack, open-source RISC-V GPGPU that supports OpenCL and AMD’s HIP programming model through chipStar and can scale to 32 cores on an Altera Stratix 10 FPGA.
Vortex’s HIP support allows kernels and serving-layer code developed in this project to use the same programming model as the companion project on AMD Instinct GPUs. Fairness mechanisms can therefore be transferred between the soft GPU and commercial AMD datacenter hardware. Although ROCm provides an open runtime, commercial GPUs ultimately conceal their hardware schedulers and memory arbiters in silicon. A soft GPU removes this barrier entirely, making Vortex a uniquely transparent platform for studying—and ultimately enforcing in hardware—fairness among multiple tenants.
The project will proceed in three stages.
First, the student will deploy Vortex on our Stratix 10 FPGAs using the Altera/OPAE toolchain. The student will validate the deployment using the standard OpenCL test suite and characterize the achievable core count, clock frequency, and memory bandwidth.
Second, the student will implement a transformer-inference workload on the soft GPU. This stage will focus on quantized inference kernels—including matrix multiplication, attention, and layer normalization—for a small language model sized to fit within the FPGA’s computation and memory constraints.
Third, the student will build a host-side serving layer that multiplexes inference requests from multiple users on the shared Vortex device. The serving layer will implement a credit-fair scheduling policy that tracks each user’s lending and borrowing of GPU time across periods of varying demand. The policy will be evaluated against FIFO and round-robin baselines using fairness, tail latency, and utilization as the primary metrics.
The transparent RTL implementation will also allow the student to instrument actual per-kernel resource usage, providing a level of accounting accuracy that is not possible on commercial GPUs. The final report will document the deployment, characterize the measured fairness–utilization tradeoff, and identify opportunities for moving fair arbitration from the host software into the GPU microarchitecture—the longer-term research direction enabled by this testbed.
You will learn about: open-source GPU architecture and RISC-V GPGPUs; FPGA deployment using Quartus and OPAE on datacenter-class hardware; GPU kernel programming using AMD’s HIP model and OpenCL; and fair scheduling mechanisms for shared accelerators.
Required skills: Strong C/C++ and Linux systems-programming skills; FPGA and RTL experience using Verilog or SystemVerilog; and familiarity with GPU programming or computer architecture. Experience with HIP/ROCm or OpenCL is an asset.
Supervisor: Prof. Seyed Majid Zahedi
Email: smzahedi@uwaterloo.ca
Linux Perf Integration for Hardware Performance Counters on a RISC-V CVA6 Processor
Hardware performance counters (HPCs) expose low-level microarchitectural events — cache misses, branch mispredictions, pipeline stalls, and more; that are essential for application profiling and optimization. The RISC-V CVA6 processor defines a rich set of such counters, but Linux perf support for them remains limited, restricting developers from using standard profiling workflows on this platform. This project will implement full Linux perf support for the CVA6 HPCs by extending the RISC-V perf kernel driver, adding the necessary OpenSBI HPM extensions for privileged counter access, and mapping CVA6-specific hardware events to the perf event model. The student will validate the implementation by profiling real workloads on an FPGA-deployed CVA6 based multicore SoC and measuring monitoring overhead.
Required knowledge & skills: Linux kernel programming (C), computer architecture fundamentals, familiarity with the RISC-V privileged ISA. Prior exposure to embedded Linux development is appreciated but not required.
Supervisor: Prof. Rodolfo Pellizzoni
Email: rpellizz@uwaterloo.ca
Location: E5 4113
LLM-Assisted Production Intelligence and Closed-Loop Control via Digital Twin Environments
This project develops an LLM-based intelligent assistant integrated with a virtual digital twin to streamline industrial automation control. By combining synthetic operational data streams with technical documentation and SOPs, the system analyzes trends, diagnoses anomalies, and executes closed-loop parameter adjustments to significantly reduce operator cognitive load and optimize overall manufacturing performance.
Required knowledge & skills: Familiarity with large language models and prompt engineering; experience with retrieval-augmented generation (RAG) and document-grounded QA systems; understanding of automation systems, PLC/SCADA concepts, and production workflows; proficiency in Python; experience with data analysis and time-series production data; knowledge of API integration and orchestration frameworks (e.g., LangChain, LlamaIndex, or similar); familiarity with digital twin platforms is an asset.
Application Instructions: Interested applicants should combine their CV, undergraduate transcripts, and graduate transcripts into a single PDF file and email it directly to the supervisor. This project is conducted in collaboration with an industry partner, offering strong potential for future co-op opportunities.
Supervisor: Professor Ladan Tahvildari
Path Planning/Controls for Autonomous Racing
In this project you will be designing, implementing, validating, and iterating upon various path planning and control algorithms to drive a modified Dallara IL-15 Indy Lights vehicle around the Indianapolis Motor Speedway in simulation. This project is a part of Waterloo Autonomous Racing (WATORACE)’s stack for competing in the Indy Autonomous Challenge.
Supervisor: Prof. Derek Rayside
Email: drayside@uwaterloo.ca
Performance Comparison of Rule and Integrity Checkers
Modern safety-critical systems require runtime monitoring to ensure integrity and safety. At the same time, these systems remain energy efficient to support small device size and operate without fans. The goal of this project is to evaluate runtime monitoring frameworks and perform a gap analysis which can then lead to subsequent research.
You will learn about: runtime verification, stream processing, embedded software, safety-critical systems, data analysis, performance evaluation
Supervisor: Prof. Sebastian Fischmeister
Email: sebastian.fischmeister@uwaterloo.ca
Prompt Optimization for Agentic LLM Systems
Automated prompt optimization has proven effective for individual LLM agents through evolutionary frameworks such as GEPA. However, these approaches encounter fundamental challenges when multiple interacting agents are optimized jointly. First, numerical optimizers cannot easily determine which agent’s prompt caused a collective failure, creating a credit-assignment problem. Second, when the prompts of interacting agents are modified concurrently, the environment faced by each agent changes continuously, creating non-stationarity and making evaluations unreliable.
Our group has developed scheduling techniques—including LLM-based semantic credit assignment and sequential, agent-by-agent optimization—that extend GEPA to multi-agent settings. These techniques have produced relative performance improvements of up to 32% in coordination-intensive games.
In this project, the student will extend this research in one of two directions, selected in consultation with the supervisor:
1. New cooperative and competitive benchmark environments: The student will develop environments that vary the game-theoretic structure of the task, including simultaneous versus sequential actions and tightly versus loosely coupled interactions. These environments will be used to test our finding that the order in which agents are optimized should reflect the task’s backward-induction structure.
2. Reward-machine-based credit assignment: The student will represent temporally extended task requirements as finite-state reward structures that help localize failures across both agents and time.
The principal deliverable will be a working extension of our optimization framework, accompanied by a rigorous empirical comparison with the scheduling techniques currently implemented in the system.
You will learn about: multi-agent systems and game theory; LLM-based agents and automated prompt optimization; and experimental design for stochastic learning systems.
Required skills: Strong Python skills; experience with LLM APIs or agentic frameworks; and a background in machine learning. Experience with multi-agent reinforcement learning or game theory is an asset.
Supervisor: Prof. Seyed Majid Zahedi
Email: smzahedi@uwaterloo.ca
Pwn-a-Truck: Cybersecurity of Heavy Vehicles
Security of autonomous vehicles is crucial to eventually deploy them at scale. We own a truck that we use for cybersecurity. The goal of the project is to identify exploitable vulnerabilities in electronic control units of an actual truck on campus. Pwn a truck!
You will learn: embedded systems security, low level programming, CAN, cybersecurity attack tools
Supervisor: Prof. Sebastian Fischmeister
Email: sebastian.fischmeister@uwaterloo.ca
Root-Cause Analysis for Safety and Security Incidents
Security and safety are paramount for modern systems like autonomous vehicles, airplanes, and medical devices. The challenge is to reason about incidents in such systems. The goal of the project is to review open-source reasoning frameworks and build a prototype for incident response for embedded systems.
You will learn about: root-cause analysis, data analysis, reasoning and AI, embedded systems, safety-critical systems
Supervisor: Prof. Sebastian Fischmeister
Email: sebastian.fischmeister@uwaterloo.ca
Scalable Deep Learning
The fast-growing size of deep learning models and datasets used for model training have accelerated the demand for novel scalable computing techniques that efficiently utilize high-performance, however, expensive, computing resources. In this project, we will investigate practical scalable computing solutions for problems in deep learning, e.g., training CNNs, LLMs or GNNs on large datasets, neural architecture search, retrieval augmented generation and federated learning. For solution design, we will explore techniques such as data and model -parallel training, pipeline parallelism, heterogenous memory for weight offloading in large models, low-bit quantization and low-rank adaptation.
Required knowledge & skills: Object-oriented programming, algorithms and data structures, familiarity with parallel/distributed computing concepts and implementation frameworks (e.g., CUDA, MPI), knowledge of machine learning concepts and tools (e.g., PyTorch, DeepSpeed).
Supervisor: Prof. Tahsin Reza
Email: tahsin.reza@uwaterloo.ca
Sim-to-Real Applications for Autonomous Driving
As driving simulators become more photorealistic and simulation platforms become more flexible in terms of sensor configuration and environmental setting, synthetic data has a greater chance of filling gaps in existing real-world datasets. Nonetheless, there is a domain difference between the appearance of simulation and real-world data, which can be bridged using transfer learning techniques. Watonomous researchers have been working on developing such schemes for autonomous driving applications. Candidates with programming experience with PyTorch/ TensorFlow and familiarity with the CARLA simulator are preferable.
Supervisor: Prof. Derek Rayside
Email: drayside@uwaterloo.ca
Snapshot-Based Performance Counter Access for SoC Monitoring
Modern system-on-chip platforms rely on hardware performance counters to monitor system activity and support runtime resource management. In the CVA6-based platform used in this project, the Advanced Platform Monitoring Unit (APMU) provides a set of hardware counters that monitor system-level events such as cache misses, interconnect traffic, and memory requests. These counters are accessed by software running on an embedded Ibex core within the APMU. However, because counters continuously increment during execution, software reads may observe inconsistent values when multiple counters are accessed sequentially. This project aims to design a snapshot mechanism that allows software to capture a consistent view of all performance counters at a specific point in time. The solution will involve adding a shadow register structure that stores counter values when triggered by a special instruction or control signal from the embedded Ibex control core. The student will implement the snapshot logic in RTL and integrate it into the existing monitoring infrastructure.
Required knowledge & skills: digital logic design, RTL (Verilog/SystemVerilog), computer architecture fundamentals. Familiarity with RISC-V architecture is appreciated but not required.
Supervisor: Prof. Rodolfo Pellizzoni
Email: rpellizz@uwaterloo.ca
Location: E5 4113
Software implementation of FRI protocol in zkSNARK systems
The blockchain privacy is implemented by a zero knowledge succinct noninteractive argument of knowledge (zkSNARK) proof system. FRI (fast Reed Solomon code Proximity) protocol is a popular protocol employed in a number of efficient and practical zkSNARK systems. The project is to implement this protocol and test the performance when it is embedded into the existing zero knowledge proof systems for blockchain privacy.
Supervisor: Prof. Guang Gong
Email: ggong@uwaterloo.ca
Synthetic Data Generation and Digital Twin Development for Visual Neural Network Training
Modern automation systems increasingly rely on computer vision neural networks to perceive and interpret their environments. However, acquiring sufficient real-world labeled data to train these models is costly, time-consuming, and often bottlenecked by safety or accessibility constraints. To address these challenges, this project explores advanced synthetic data generation techniques tailored for industrial automation applications. Specifically, the project investigates the creation of high-fidelity synthetic datasets that accurately reflect complex automation environments. Furthermore, the project will develop a digital twin of the target automation application. This digital twin will serve as a high-fidelity, virtual testbed, enabling the trained model to be safely deployed and validated in a controlled environment that mirrors the physical system. By decoupling development from physical hardware, this approach facilitates iterative model improvement and rapid prototyping of visual perception pipelines.
Required Knowledge & Skills: Knowledge of deep learning fundamentals; experience with synthetic data generation tools and simulation environments; familiarity with digital twin concepts and 3D modeling; proficiency in Python and relevant ML frameworks; knowledge of image annotation and dataset management workflows.
Application Instructions: Interested applicants should combine their CV, undergraduate transcripts, and graduate transcripts into a single PDF file and email it directly to the supervisor. This project is conducted in collaboration with an industry partner, offering strong potential for future co-op opportunities.
Supervisor: Professor Ladan Tahvildari
Urban Decision Making for Autonomous Vehicles
Developing decision-making frameworks for safe, yet efficient, urban autonomous driving under environment uncertainties is a challenging topic for modern learning-based algorithms. Therefore, in this project, reinforcement learning-based decision-making schemes are developed to learn optimal policies for driving in multi-agent environments where intents of road users and other dynamic states are unforeseeable. Candidates with programming experience with PyTorch/ TensorFlow and familiarity with the CARLA simulator are preferable.
Supervisor: Prof. Derek Rayside
Email: drayside@uwaterloo.ca
Urban Decision Making for Autonomous Vehicles
Developing decision-making frameworks for safe, yet efficient, urban autonomous driving under environment uncertainties is a challenging topic for modern learning-based algorithms. Therefore, in this project, reinforcement learning-based decision-making schemes are developed to learn optimal policies for driving in multi-agent environments where intents of road users and other dynamic states are unforeseeable. Candidates with programming experience with PyTorch/ TensorFlow and familiarity with the CARLA simulator are preferable.
Supervisor: Prof. Derek Rayside
Email: drayside@uwaterloo.ca