Industry

Senior Engineer, ArchitectureTenstorrent USA, Inc., Santa Clara, CA
September 2025 – present

Performance Modeling and Characterization of Agentic AI on Emerging CPU Architectures with Custom Accelerators (April 2026 – present)

  • Characterize performance bottlenecks of representative agentic AI pipelines (large language model inference, retrieval, planning, and action execution) on heterogeneous platforms combining CPUs, GPUs, and custom AI accelerators.
  • Build predictive performance models that capture compute, memory, and communication behavior across devices, including microarchitectural profiling of irregular branching and cache behavior.
  • Develop hardware–software co-optimizations, including workload partitioning, memory-layout transformations, scheduling strategies, accelerator-specific kernels, and memory prefetching, to reduce end-to-end latency and energy consumption.

Hardware-Algorithm Co-optimization to Accelerate Server-Class Machine Learning Workloads on Emerging RISC-V CPUs (September 2025 – present)

  • Identify performance bottlenecks and accelerate server-class machine learning workloads on RISC-V CPU platforms by optimizing data movement, memory hierarchy utilization, and workload mapping across cores and vector units.
  • Collaborate with processor architecture, performance analysis, and compiler optimization teams to guide microarchitectural decisions, including vector unit configurations, cache hierarchy tuning, and memory subsystem design.
  • Develop vectorization strategies and multi-core workload-distribution methods that exploit spatial and temporal data locality to reduce memory bandwidth pressure.

Engineer – Wave Computing, Santa Clara, CA
February 2018 – July 2019

Research Intern – DEVCOM Army Research Lab (DEVCOM ARO), Marina del Rey, CA
May 2022 – August 2022

Academic

Research Assistant – Ming Hsieh Department of Electrical and Computer Engineering, USC
August 2019 – August 2025

Research highlights (supervised by Prof. Viktor Prasanna):

  • Tensor decomposition on CPU, GPU, and FPGA: Parallel algorithms and novel tensor formats to reduce sparse MTTKRP time; custom designs across CPU, GPU, and FPGA.
  • Photonic SRAM / SPRINT: Hardware-algorithm co-design mapping MTTKRP to SPRINT; performance models for peak and sustained performance; collaboration with the hardware team on bottlenecks.
  • GNN for SAR-ATR: Explainable models for GNNs on SAR imagery; human-in-the-loop to improve accuracy.

Shared memory controller (2020-2021): Memory controller for heterogeneous platforms handling irregular traffic from CPU, GPU, and FPGA via routing/mapping.

Prior (University of Moratuwa): Reconfigurable co-processor with limited precision for DNNs; SDN switch on FPGA (OpenFlow); HEVC SCC extension (Intra Block Copy, Palette Coding) at ParaQum Technologies.