HKContact
§ 07 — Contact

Open to research, residencies, and roles in LLM systems and foundation models.

Research Interests

Inference runtime optimization, Mixture-of-Experts (MoE) caching, hardware-software co-design, model compression, and post-training data alignment.

Currently Interested In

Low-latency MoE offloading hierarchies, kernel optimization for attention scaling, and fact-grounded SFT data synthesis.

Open To

AI research engineer roles, systems residency programs, and open research collaborations on efficient models training.

Availability

Immediate (Remote / Hybrid / On-site).

© 2026 Harsh Kaushik. All rights reserved.
Currently: TurboLLM · inference research