Keisuke Kamahori (釜堀 恵輔; he/him)
kamahori [at] uw.edu
I'm a third-year Ph.D. student at Paul G. Allen School of Computer Science & Engineering, University of Washington, advised by Baris Kasikci at SyFI Lab.
I'm working on the following topics:
-
Agents & Systems:
I'm interested in how AI agents are reshaping the way we build systems. VibeServe★ is a multi-agent system that synthesizes LLM serving systems end-to-end, rethinking how ML infrastructure should be built in the era of AI agents. AgentFlux is an efficient fine-tuning and inference framework for local AI agents with the idea of decoupled post-training.
I'm a part of the Self-Defining Systems project. I also won multiple awards at the MLSys 2026 Competition (NVIDIA Track) on agentic CUDA generation. -
Systems for Multimodal Models:
I see multimodal models as the next frontier for inference optimization, with serving demands that don't fit cleanly into existing stacks. M* is a universal system for serving any multimodal models with computation graph abstraction. Murmur is an efficient inference system for long-form ASR. VoxServe★ is a streaming-centric serving system for speech language models, designed around the unique latency and continuity demands of speech applications. LiteASR★ (EMNLP 2025 Main) is a compression technique for speech models based on low-rank approximation.
I worked as a Founding Engineer at Kotoba Technologies, where I built the serving system behind the world's fastest simultaneous translation app. I also collaborate with AI4DeafBlind.org to enable fast ASR inference on low-resource accessibility devices. -
Local AI:
TeleRAG★ (MLSys 2026) is an inference acceleration system for RAG that hides retrieval latency via lookahead retrieval, and Fiddler★ (ICLR 2025) accelerates Mixture-of-Experts inference under tight resource constraints through CPU-GPU orchestration. ConsumerBench (COLM 2026) is a benchmarking framework for generative AI applications running on end-user devices. -
Computer Architecture:
Beyond ML systems, I've always been excited about computer architecture and hardware acceleration. I will join NVIDIA as an Architect Intern in Summer 2026. Themis★ is a software-defined hardware prefetching mechanism for data center CPUs. Prior to my PhD, I worked extensively on computer architecture, including Manticore (ASPLOS 2024, FPGA 2024), a hardware-accelerated RTL simulator.
★ projects I led / co-led.
I'm honored to be a recipient of the Toyota Riken Ph.D. Scholarship (2023-2025).
Prior to that, I received B. Sc. in Information Science from the University of Tokyo in 2023, advised by Shinya Takamaeda-Yamazaki. I also worked with James Larus at EPFL in the summer of 2022.
[CV] [Google Scholar] [ORCID] [DBLP] [Linkedin] [X]
Preprints
-
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
Keisuke Kamahori*, Shihang Li*, Simon Peter, Baris Kasikci
arXiv preprint; DL4C workshop @ ICML 2026
Blog Paper … -
M*: A Modular, Extensible, Serving System for Multimodal Models
Atindra Jha*, Naomi Sagan*, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang
arXiv preprint
Project Page Paper … -
MURMUR: An Efficient Inference System for Long-Form ASR
Wei-Tzu Lee, Keisuke Kamahori, Baris Kasikci
arXiv preprint
Paper … -
VoxServe: Streaming-Centric Serving System for Speech Language Models
Keisuke Kamahori, Wei-Tzu Lee, Atindra Jha, Rohan Kadekodi, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci
arXiv preprint
Blog Paper … -
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
Rohan Kadekodi*, Zhan Jin*, Keisuke Kamahori, Yile Gu, Sean Khatiri, Noah H. Bayindirli, Sergey Gorbunov, Baris Kasikci
arXiv preprint
Project Page Paper … -
Themis: Software-Defined Hardware Prefetching
Keisuke Kamahori, Neil Adit, Kan Zhu, Yuqi Mai, Victor Lee, Heiner Litz, Chris Kennelly, Snehasish Kumar, Hanna Alam, Milad Hashemi, David Li, Adrian Sampson, Baris Kasikci, Tipp Moseley, Parthasarathy Ranganathan, Akanksha Jain
arXiv preprint
Paper
Peer-Reviewed Publications
-
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
Yile Gu*, Rohan Kadekodi*, Hoang Nguyen, Keisuke Kamahori, Yiyu Liu, Baris Kasikci
COLM 2026
Project Page Paper … -
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
Chien-Yu Lin*, Keisuke Kamahori*, Yiyu Liu*, Xiaoxiang Shi, Madhav Kashyap, Yile Gu, Rulin Shao, Zihao Ye, Kan Zhu, Rohan Kadekodi, Stephanie Wang, Arvind Krishnamurthy, Luis Ceze, Baris Kasikci
MLSys 2026
Paper … -
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
Keisuke Kamahori, Jungo Kasai, Noriyuki Kojima, Baris Kasikci
EMNLP 2025 Main
Paper …Models
-
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
Keisuke Kamahori*, Tian Tang*, Yile Gu, Kan Zhu, Baris Kasikci
ICLR 2025; PML4LRS workshop @ ICLR2024
Paper … -
NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Tian Tang, Qinyu Xu, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci
OSDI 2025
Paper … -
A 475 MHz FPGA Accelerator for RTL Simulation
Sahand Kashani*, Mahyar Emami*, Keisuke Kamahori, Mohammad Sepehr Pourghannad, Ritik Raj, James R Larus
FPGA 2024
DOI Code -
Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
Mahyar Emami*, Sahand Kashani*, Keisuke Kamahori, Mohammad Sepehr Pourghannad, Ritik Raj, James R Larus
ASPLOS 2024
Paper DOI Code -
CiraaS: cloud computing with programmable logic
Kenji Tanaka, Yuki Arikawa, Tsuyoshi Ito, Yuki Matsuda, Keisuke Kamahori, Shinya Kaji, Takeshi Sakamoto
SIGCOMM 2022 Poster
DOI -
Accelerating Decision Tree Ensemble with Guided Branch Approximation
Keisuke Kamahori, Shinya Takamaeda-Yamazaki
HEART 2022
DOI