The Hybrid Observability Problem, Solved with Selector


This video is part of the appearance, "Selector AI Presents at Cloud Field Day 26". It was recorded as part of Cloud Field Day 26 at 10:30am - 12:00pm on September 23, 2026.


Watch on YouTube
Watch on Vimeo

Reza Koohrangpour of Selector framed hybrid cloud observability as an operational problem more than a technological one. Each cloud provider, whether AWS, Azure, or Google Cloud, brings its own services, data models, and telemetry, while the enterprise side adds SNMP and other traditional networking protocols. These domains have operated as separate silos for years, so when an incident occurs, operators must manually correlate data across them and speak every one of those languages. Koohrangpour also pointed to the cost side: every new tool adds another console, another skill set, and another integration that someone has to build and maintain by hand.

Koohrangpour argued that the largest and least discussed cost in mean time to resolution is coordination, not diagnosis. Because incidents span so many domains, teams spend time pulling in specialists, asking the cloud engineer to join the bridge and bring evidence, and working out what the problem is and who owns which piece. That effort never shows up on a dashboard. He gave three reasons the problem persists. First, application paths constantly shift across data centers, cloud, and on-premises networks, so there is no reliable dependency map, and the topology an operator sees is often a week old and unrelated to the current incident. Second, there are many potential failure points, including accounts, service accounts, regions, VPCs, and transit, with the gaps between them, such as transit gateways and interconnects, being the hardest to troubleshoot because nobody clearly owns them. Third, providers lack a standard data model.

He then explained why existing tools have not closed this gap, attributing it to the lack of a unifying standard in three areas. Constructs differ across platforms, so an operator investigating a single Kubernetes incident may need to navigate GKE, AKS, EKS, OpenShift, and Nutanix terminology. Data models also diverge: Azure organizes around resource groups, AWS around accounts, and Google Cloud around projects, making even a simple operational question hard to answer consistently. Logs add further friction, arriving as syslog, pub/sub messages, or query results, in varying versions, formats, and locations, which slows any investigation. The transcript ends after this problem statement, before Koohrangpour describes how Selector’s platform addresses these challenges.

Personnel:

Sign up for updates to
Tech Field day events

Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.

We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.

Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!