Watch on YouTube
Watch on Vimeo
John Heintz, who leads Selector’s pre-sales engineering team, demonstrates how an operator would use the platform to troubleshoot application reachability problems whose causes are hidden across hybrid infrastructure. In the first scenario, an operator named Michael receives a Slack alert (ServiceNow or another ticketing tool would work the same way) reporting that an application called Cloud Field Day is unreachable from one probe, likely because of firewall drops at the VPC level, while the interconnects appear healthy. Heintz clarified that Selector itself, not the LLM, identifies the cause by detecting anomalies such as firewall drop rates deviating from their baseline, and the LLM only renders the collected data into readable text and recommendations, optionally drawing on company runbooks. Following the alert’s link into the portal, the operator sees that the VM is up, that Selector’s PingMesh probe shows the application unreachable, and that a firewall rule edit denying traffic from an IP range sits on the path between probe and application. Anomaly thresholds can be static values or machine learning baselines tuned per metric, with ML favored for drops, utilization, errors, and discards. Heintz noted that Rosetta routes requests to specialized agents, so keywords like “debug” or “root cause” prompt the Loom investigation agent to run a full analysis while simpler questions get direct answers.
Heintz then showed the same incident investigated conversationally, asking Rosetta step by step about reachability, the cloud path, and interconnect health, with each answer drawing on different skills. Delegates asked whether the system learns from repeated investigations, and Heintz explained that users can refine answers through feedback, which builds new skills on the fly, while a thumbs-down rating and pre-deployment QA with expected responses help keep answers accurate. Conversations are stored and can be queried later, such as asking how a colleague fixed an issue last week, and multi-user conversations are in development. Before enabling agentic features, Selector builds a metadata store of customer application names, site names, locations, and terminology, continuously refreshed through APIs to tools such as ServiceNow, Dynatrace, and AppDynamics, and it can ingest Confluence pages and Visio diagrams. He pointed out where the LLM added genuine value, explaining that Google firewall rules are evaluated strictly by priority, and showed that a single generic “debug” prompt runs the entire multi-step investigation at once. Investigations rely mainly on ingested data but can use MCP to fetch information at runtime, and Rosetta can take corrective actions if granted permissions under its own service identity, though most customers keep accounts read-only and route changes through partners such as Itential with a human in the loop.
The second scenario involved an application in Chicago whose VMs were healthy but unreachable because two BGP sessions had dropped, removing the routes to the application even though the physical layer was fine. Heintz emphasized that the dashboard widgets are generated dynamically from the devices and events in each incident, and that stitching together data plane and control plane information helps separate application, infrastructure, network, and cloud teams avoid the usual all-hands bridge call. Asked whether Selector could detect provider outages before AWS acknowledges them, he said the platform ingests cloud provider health dashboards but doesn’t correlate data across customers, though a delegate’s suggestion to run Selector-owned probes across regions to create a shared signal was welcomed as a good idea. For on-premises deployments, Rosetta can use a customer’s approved LLM instead of the default Gemini by plugging in its endpoint and key, with per-customer prompt templates to handle model quirks, and it works across languages. Heintz noted that Rosetta has been available for only about six months, and the platform functions fully without it. He closed by describing how Rosetta can triage on-premises networks, cloud networks, and cloud applications in a single investigation, and how carrier data can even connect a root cause like a fiber cut to the construction crew that caused it.
Personnel: John Heintz
Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.
We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.
Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!