Tanish Samir Desai
tanishdesai37@gmail.com | Vadodara, India
LinkedIn | GitHub | Google Scholar | Resume
Research
GPU Power Cap Optimization (HPDC '25 – Best Poster Award)
View Publication
ACM HPDC '25 Symposium, Notre Dame, USA
- Developed a machine learning–driven framework that dynamically adjusts GPU power caps to maximize performance-per-watt within constrained energy budgets.
- Conducted extensive evaluation across 20+ diverse workloads from Rodinia and Parboil benchmark suites, achieving up to 18% higher computational throughput compared to NVIDIA's baseline power management strategies, while consuming equivalent average power.
- Demonstrated scalability benefits in multi-GPU cluster environments, where our approach yielded 15–20% energy savings per compute node, translating to substantial cost reductions at enterprise HPC data center scales.
- Recognized with the Best Poster Award at HPDC '25, underscoring the critical intersection of artificial intelligence and sustainable computing practices in high-performance environments.
- Created a novel static analysis–based prediction system that estimates GPU power consumption without requiring expensive hardware instrumentation or runtime profiling.
- Conducted comprehensive comparative analysis of multiple machine learning paradigms—including gradient boosting, deep neural networks, and transfer learning techniques—across six distinct CUDA benchmark collections.
- Enabled rapid power profiling by transforming traditionally time-intensive measurement workflows (hours) into near-instantaneous static analysis (seconds), empowering both kernel optimization researchers and infrastructure teams managing energy efficiency metrics.
- Engineered a sophisticated LSTM-based neural architecture capable of accurately forecasting GPU thermal states by modeling intricate temporal dependencies and workload evolution patterns.
- Leveraged recurrent learning capabilities to capture dynamic runtime characteristics, outperforming traditional static prediction approaches through adaptive, context-aware thermal forecasting.
- Delivered practical value in operational settings by enabling preemptive thermal intervention strategies, which mitigate performance throttling events, enhance system reliability margins, and prolong hardware service lifetimes in large-scale data center and HPC deployment environments.