Research Focus
On-Device ML & Edge NPUs
Accelerating inference across resource-constrained devices, spanning Transformers, CNNs, and RNNs through quantization and compiler optimizations.
- Efficient PyTorch graph & kernel implementations
- Model compression & int8/int4 quantization
- SoC vendor collaboration on NPU/eNPU specs
OS & Memory Systems
Architecting virtual memory, lazy translation coherence, and operating systems tailored for many-core and heterogeneous compute platforms.
- Lazy Translation Coherence (LATR - ASPLOS'18)
- Eventually Consistent TLBs (ECOTLB - ACM TACO'20)
- Data-centric OS for accelerators (SOLROS - EuroSys'18)
Secure & Scalable Systems
Developing high-throughput distributed graph processing and hardware enclave security mechanisms for virtualized network functions.
- Trillion-edge graph engine (Mosaic - EuroSys'17)
- SGX-secured state isolation (S-NFV - Best Paper)
- Ordered TCP server processing (TCP Ordo - INFOCOM'16)
Featured Systems & Projects
Selected open-source autonomous agents, edge AI tools, and browser science sandboxesZenith: Local-First Autonomous Coding Agent
Modular autonomous coding assistant running on-device on Apple Silicon or cloud models. Features 2-tier persistent memory, AST symbol indexing, and a resilient 5-tier fuzzy edit engine with pre-mutation /undo checkpoints.
AR Physics Lab
In-browser experimental physics powered by real-time computer vision. Augment a falling ball over live camera feeds, measure gravitational acceleration (g = 2h/t²), track colored spheres, and export experiment trials.
Stock Monitoring Dashboard
High-density tech market dashboard monitoring 60 tickers across 6 sectors (Big Tech, AI hardware, Cloud, Security). Ingests and caches live Yahoo Finance data hourly via automated GitHub Actions pipelines.
On-Device AI Image Editor
Zero-cloud photo editing using natural language prompts powered by Qwen-Image-Edit running locally via Apple Silicon MLX/mflux, featuring 4x Real-ESRGAN upscaling and persistent worker architecture.
Experience & Education
Focused on accelerating machine learning inference on embedded devices. My work encompasses optimizing and deploying diverse ML architectures—including CNNs, RCNNs, RNNs, and transformers—onto resource-constrained hardware platforms. I leverage PyTorch and related frameworks to develop efficient model implementations and quantization techniques for edge deployment. Additionally, I partner with SoC vendors to architect NPU/eNPU specifications that enable more efficient and effective hardware solutions for on-device ML inference.
Specialized in Computer Systems under the advisement of Dr. Taesoo Kim. Research focused on operating systems, virtual memory translation coherence (TLBs), low-latency data center architectures, and enclave-protected network functions. Doctoral thesis: Taming Latency In Data Center Applications.
Professional Service & Engagement
Publications
570+ Citations · Google ScholarECOTLB: Eventually Consistent TLBs
@article{maass2020ecotlb,
title={ECOTLB: Eventually Consistent TLBs},
author={Maass, Steffen and Kumar, Mohan and Kim, Taesoo and Krishna, Tushar and Bhattacharjee, Abhishek},
journal={ACM Transactions on Architecture and Code Optimization (TACO)},
volume={17},
number={4},
pages={1--25},
year={2020},
publisher={ACM}
}
SOLROS: A Data-Centric Operating System Architecture for Heterogeneous Computing
@inproceedings{min2018solros,
title={SOLROS: A data-centric operating system architecture for heterogeneous computing},
author={Min, Changwoo and Kang, Woon-Hak and Kumar, Mohan and Kashyap, Sanidhya and Maass, Steffen and Jo, Heeseung and Kim, Taesoo},
booktitle={Proceedings of the Thirteenth EuroSys Conference},
pages={1--15},
year={2018}
}
LATR: Lazy Translation Coherence
@inproceedings{kumar2018latr,
title={LATR: lazy translation coherence},
author={Kumar, Mohan and Maass, Steffen and Kashyap, Sanidhya and Vesel{\`y}, Jan and Yan, Zi and Kim, Taesoo and Bhattacharjee, Abhishek and Krishna, Tushar},
booktitle={Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)},
pages={651--664},
year={2018}
}
Mosaic: Processing a Trillion-Edge Graph on a Single Commodity Machine
@inproceedings{maass2017mosaic,
title={Mosaic: Processing a trillion-edge graph on a single commodity machine},
author={Maass, Steffen and Min, Changwoo and Kashyap, Sanidhya and Kang, Woonhak and Kumar, Mohan and Kim, Taesoo},
booktitle={Proceedings of the Twelfth European Conference on Computer Systems (EuroSys)},
pages={527--543},
year={2017},
note={Best Student Paper Award}
}
S-NFV: Securing NFV states by using SGX
@inproceedings{shih2016s,
title={S-NFV: Securing NFV states by using SGX},
author={Shih, Ming-Wei and Kumar, Mohan and Kim, Taesoo and Gavrilovska, Ada},
booktitle={Proceedings of the 2016 ACM International Workshop on Security in Software Defined Networks & Network Function Virtualization (SDN-NFV Security)},
pages={45--48},
year={2016},
note={Best Paper Award}
}
TCP Ordo: The cost of ordered processing in TCP Servers
@inproceedings{kumar2016tcp,
title={TCP ordo: The cost of ordered processing in TCP servers},
author={Kumar, Mohan and Gavrilovska, Ada},
booktitle={IEEE INFOCOM 2016 - The 35th Annual IEEE International Conference on Computer Communications},
pages={1--9},
year={2016},
organization={IEEE}
}