Watch on YouTube
Watch on Vimeo
Keyur Patel of Arrcus laid out the problem space behind the Arrcus Inference Network Fabric, explaining how agentic and physical AI applications are driving distributed inference traffic that behaves very differently from training or web traffic. He covered the latency, sovereignty, model diversity, and capacity requirements this places on the network and the operator segments a dedicated inference fabric could serve. Patel, the company’s founder and CTO, presented at Networking Field Day 41 ahead of chief architect Nalin Pai’s walkthrough of the solution itself.
Patel split inference into physical AI, such as drones, patient monitoring, and robot-assisted surgery, and agentic AI, such as applications handling 911 dispatch or drafting medical notes. He cited OpenRouter data showing year-over-year growth in token consumption and Cisco reports showing inference flows lasting longer than web flows, with upstream traffic rising as each request carries more context. KV cache transfer dominates inference traffic, and the metrics that matter are time to first token, time per output token, and p90 and p99 latency. Patel argued that as GPUs and memory shrink end-to-end token latency, network latency will become the visible bottleneck, just as it did in training. Responding to a delegate, Arrcus noted that human-facing applications prioritize time to first token while others favor throughput or the cheapest GPU, and pointed to a public safety proof of concept with TELUS in Canada as a highly latency-sensitive case.
When a delegate questioned how this differs from back-end congestion control, Arrcus answered that distributed inference shifts the problem to the front end, where the task is steering each request to the right site with the right model and capacity. Patel added data growth, geofenced sovereignty requirements, model diversity and auto-discovery, and uneven site power and capacity to that list, concluding that the answer is a holistic fabric that keeps request and response latency bounded rather than a component-level fix. He identified neoclouds steering traffic to GPUs, interconnect providers selling SLA-backed inference, sovereignty-focused enterprises, GPU-as-a-service operators, and colocation providers converting facilities into inference sites as the segments a single inference router could serve.
Personnel: Keyur Patel, Nalin Pai, Sanjay Kumar
Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.
We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.
Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!