Skip to content

About

I couldn’t leave the problem alone.

Computer Engineering at SCET, Surat. B.Tech, 2022–2026, 9.94 CGPA. But the number I’m actually proud of is smaller, and measured in milliseconds.

Voice interfaces fascinated me because they were the ones that kept failing. Type a query and wait, fine. But talk to a machine and wait, and the whole illusion breaks. The pause tells you the truth: nobody’s really listening.

I wanted to close that gap, not fake it with a spinner, actually close it. That obsession turned into a job: today I do MLOps at VideoSDK, where “the model works” is the starting line, not the finish.

// the part nobody sees: the 200 milliseconds

Meet Jariwala
Meet Jariwala
“A demo forgives latency. A conversation doesn’t.”

I make voice AI fast enough to feel human, and reliable enough to run in production.

10–50ms

The inference server

A gRPC inference server in Rust, serving at P99, the number that decides whether a conversation feels alive or dead.

plain: how fast the AI answers you.

21,584 hrs

The speech pipeline

A training pipeline on AWS S3, 12.9M samples across 13 languages. Multilingual is the foundation, not an afterthought.

plain: everything it learned to listen from.

+34 pts

The turn-detector

A model that knows the instant you're done speaking, beating LiveKit's by 34 points. Getting it right is invisible; that's the point.

plain: it knows when to stop and reply.

−80%

The efficiency win

An AMD-optimized model that cut token usage by 80%. Same quality, a fifth of the cost to run.

plain: same brain, five times cheaper.

Builder. Creator. Speaker.

These aren’t three jobs. They’re one loop.

I build

the systems, the servers, the pipelines, the models that hold up when real traffic hits them.

I create

the artifacts around them, the writing, the docs, the demos that make invisible engineering legible.

I speak

I get in the room and explain the hard thing simply. A system nobody understands is a system nobody trusts.

“I build the systems, then I explain them to rooms. Same job.”

Numbers, not adjectives.

10–50ms

P99 inference
latency

This is how long you wait to be understood. Everything below is in service of this number staying small.

Award

Best of CES 2026

Recognized at the largest tech stage in the world, for a product, not a pitch deck.

Publication

Published in IEEE Xplore

My work is in the permanent technical record. Peer-reviewed, cited, real.

Selection · Top 0.1%

Nas Daily shortlist, Dubai

Selected from 100,000+ applicants. The half of the loop that isn't code: holding a room.

Academic

CGPA 9.94 / 10

Four years, consistently. I don't ship sloppy and I don't study sloppy.

Leadership

IoT Club Lead · Speaker

Led the team, ran the builds, taught the juniors. If I built it, I can explain it.

The honest reason, not the polished one.

The moment a machine understands you instantly, no lag, no “processing,” no waiting, something shifts. It stops feeling like a tool and starts feeling like it’s paying attention.

That moment only exists on the right side of a few hundred milliseconds. Miss it and you’ve built a very expensive form.

That’s why I don’t build demos. A demo is a promise you make in a controlled room. Production is that promise kept at 3am, under load, in a language the marketing deck forgot about. I’d rather ship the second one.

“Anyone can make it work once. I make it work at 3am.”

The hardest version of this problem.

I’m heading into real-time, multilingual, multimodal AI that runs cheaply enough to be everywhere, not just in a well-funded demo.

Faster inference.

Smaller models that don't compromise.

Voice in the languages the industry ignores.

Now · Jul 2026

Shipping the Rust inference server. Deep in CUDA kernels and quantization. Building from Surat, India, open to what’s next.

Building something that has to be fast, has to be understood, and actually has to ship? That’s the entire job.

Meet Jariwala, builds the systems, then explains them to rooms.