Watch on YouTube
Watch on Vimeo
Akshay Joshi and Jeff Cassiana introduce Clumio as a SaaS-first cyber-resilience solution that provides cloud data protection and in-place recovery for modern data pipelines. Founded in 2017 and acquired by Commvault in 2024, Clumio has expanded its cloud-native support across workloads like Amazon S3, DynamoDB, RDS, Apache Iceberg, and Google Cloud Storage. The presenters address critical cloud challenges, such as dynamic infrastructure, advanced cyber threats, integration demands, and emerging risks from agentic frontier AI models, by highlighting Clumio’s serverless architecture, immutable air-gapped backups via Secure Vault, and granular in-place recovery capabilities.
To demonstrate cloud application resilience, the presentation outlines three core requirements: protecting the entire data pipeline, delivering near real-time recovery, and scaling to meet large cloud demands. These principles are applied to “Clumio Final Cut,” a simulated movie application powered by DynamoDB, Amazon S3, and Apache Iceberg, featuring user watch lists and a generative AI chatbot. The first failure scenario examines a bad canary code deployment that corrupts DynamoDB tenant partitions at different point-in-time timestamps, forcing Site Reliability Engineers (SREs) to manage partition recovery, application reconfiguration, operational costs, and downtime.
The presentation contrasts the traditional “recovery hell”, which requires doing full table restores for every impacted partition, cherry-picking data into the source table, and cleaning up temporary tables, with the Clumio Backtrack approach. Clumio Backtrack for DynamoDB offers a simplified two-step, point-in-time, in-place recovery process directly to the source table without requiring temporary resources or application reconfigurations. The session concludes with a demonstration where a simulated corruption wipes customer watch list data from a DynamoDB table, followed by a fast, in-place restoration of the records through Clumio’s Secure Vault portal.
Personnel: Jeff Cassianna
Watch on YouTube
Watch on Vimeo
Cloud-native applications depend on data distributed across databases, object stores, retrieval pipelines, and lakehouse tables. When that data is deleted, corrupted, or changed by a bad deployment, recovery can require full restores, temporary resources, application reconfiguration, and rebuilding downstream dependencies. In this Cloud Field Day 26 session, Jeff Cassianna and his colleague Akshay used a fictitious streaming app called Clumio Final Cut to show what that looks like for an AI feature. Its “AI Movie Assistant” concierge chatbot is a typical LLM application backed by a retrieval-augmented generation (RAG) pipeline with three layers: movie metadata stored as JSON, text, and PDF objects in an Amazon S3 bucket, vector embeddings of that data held in S3 Vectors, and the LLM that queries the vector store to answer questions. When the underlying S3 objects are deleted or corrupted, the embeddings point to nothing. The chatbot then either fails or starts hallucinating, for example, claiming that Love Actually is an action movie and failing to find the film “Clumio to the Moon.”
Akshay framed recovery around the SRE’s priorities. The team needs a recovery point with minimal data loss, as little LLM reconfiguration as possible (a slow and error-prone process), controlled costs, and minimal downtime. With traditional all-or-nothing tools, recovery takes five steps: restore the entire bucket to a new location, recompute the vectors, delete the old vector store, clean up the old bucket, and repoint the LLM. Clumio Backtrack for S3, launched at re:Invent 2024 after the Commvault acquisition, reduces this to selecting a point in time and the affected objects, then clicking restore. Backtrack combines three capabilities. It offers near-instant point-in-time rollback down to the second, granular recovery at the object, prefix, or bucket level, and in-place recovery, backed by continuous change tracking of hundreds of millions of object changes per hour. Responding to delegate questions, Akshay said in-place recovery suits internal incidents such as a bad code push rather than malicious attacks. He also said granular recovery effectively merges good data with restored data, and Clumio does not currently detect when an event occurred, so customers must bring that timestamp themselves. He argued that backup is now commoditized and that Clumio is investing in recovery options.
Jeff’s demo corrupted the “Clumio to the Moon” JSON object with a Python script, showed the chatbot returning inaccurate answers, and then used Backtrack to locate and overwrite the object with a prior version in under a minute, after which the chatbot returned accurate responses. Jeff noted that Backtrack moves no data; it makes S3 versioning usable at the scale of billions of objects. Secure Vault is Clumio’s isolated, air-gapped offering that copies data outside the customer’s security sphere. Backtrack requires versioning when used without Secure Vault. You can run bulk recoveries of thousands or billions of objects through the API, and Clumio has been tested with tens of billions of S3 objects and hundreds of billions of DynamoDB items. Its serverless, Lambda-based architecture is designed to avoid throttling. Restores can target a different onboarded AWS account or region, which Jeff recommended for cyber events or regional outages like the recent us-east-1 incident and which also helps with data sovereignty requirements. He closed by noting that Clumio recently achieved FedRAMP Ready status, with further compliance details at trust.commvault.com.
Personnel: Akshay Joshi, Jeff Cassianna
Watch on YouTube
Watch on Vimeo
Cloud-native applications depend on data distributed across databases, object stores, retrieval pipelines, and lakehouse tables. When that data is deleted, corrupted, or changed by a bad deployment, recovery can require full restores, temporary resources, application reconfiguration, and rebuilding downstream dependencies. In the final scenario of their Cloud Field Day 26 session, Jeff Cassianna and his colleague Akshay turned to the Box Office Analyzer feature of the fictitious Clumio Final Cut app. The feature draws on a rich dataset stored in a lakehouse built on Apache Iceberg, which Akshay described as the preferred format for analytics at scale. The data lives in Amazon S3 Tables, is queried through Amazon Athena, and is visualized in QuickSight; it serves end users, internal employees, and production houses who build their own dashboards on it. Iceberg’s weakness is that a schema change or dropped column creates a mismatch with what visualizations expect. In the example, someone altered the table’s column structure, which broke both the in-app charts and any external dashboards. The revenue-by-genre chart then showed every movie as a comedy.
From an SRE perspective, Akshay noted that the last known good recovery point for an Iceberg table can be a point in time or a snapshot, depending on whether you ask a data admin or a storage admin. An out-of-place recovery would also force teams to recreate every dashboard built on the data. Before native support existed, Iceberg data could only be protected by backing up the general-purpose S3 bucket underneath it. Recovery that way means restoring every bucket in full, rebuilding the Iceberg structure of metadata, manifest, and data files, repointing the application to the new tables, and recreating all dashboards inside and outside the app, which Akshay called “recovery hell.” Clumio, which he said was the first offering to support Iceberg, makes backups and restores that are Iceberg-aware and transactionally consistent. That preserves the table structure and reduces recovery to a table-level, in-place rollback with no reconfiguration. Clumio also places no limit on retained snapshots. Customers facing FINRA or HIPAA retention requirements can therefore keep a lean primary lakehouse and offload older snapshots, since a growing snapshot tree degrades performance. Clumio can also migrate Iceberg tables from Glue Data Catalog or general-purpose buckets to S3 Tables, which Jeff noted is only the third new bucket type in AWS’s 20-year history. The capability currently covers AWS only.
Jeff’s demo showed the AI Movie Assistant correctly identifying “Clumio to the Moon” as an action-adventure film. He then staged a double failure by corrupting both the Iceberg table and the movie’s S3 JSON object, after which the revenue chart attributed about $400 million entirely to comedy and the chatbot mislabeled the genre. He restored the table from the isolated Secure Vault copy by selecting a point in time matched against available Iceberg snapshots, and he restored the object through Backtrack for S3. Both tasks finished in about a minute, and the chart and chatbot answers returned to normal. In response to delegate questions, Akshay explained that restores work with native Iceberg snapshots. Customers can recover one, several, or the full snapshot chain and set any of them as the latest. A restore to a new table rewrites the data, while an in-place restore places the snapshot where it belongs. Jeff closed with Commvault’s guidance on architecting for resiliency in AI applications. He urged teams to protect the entire data pipeline, not isolated pieces, as agentic AI adds complexity; plan for the near-real-time resolution customers demand to avoid brand damage and churn; and choose a solution that can keep pace as apps scale to millions of users and new geographies.
Personnel: Akshay Joshi, Jeff Cassianna
Watch on YouTube
Watch on Vimeo
Cloud-native applications depend on data distributed across databases, object stores, retrieval pipelines, and lakehouse tables. When that data is deleted, corrupted, or changed by a bad deployment, recovery can require full restores, temporary resources, application reconfiguration, and the reconstruction of downstream dependencies. In the closing segment of Commvault’s Cloud Field Day 26 session, Jeff Cassianna and his colleague Akshay turned from Clumio’s AWS capabilities to its expansion across clouds. Akshay observed that applications increasingly ignore cloud boundaries, with customers assembling the best services from AWS, Google Cloud, and Azure, so Clumio is becoming a multi-cloud data protection offering. Clumio announced support for Google Cloud Storage at Google Cloud Next in May 2026, with general availability in summer. It brings the same operating principles Clumio uses for S3 backups that are air-gapped and immutable in a separate Google Cloud project, an agentless and serverless design that needs only a role with the right permissions, and scalability to the 70 to 80 billion objects per bucket already proven on S3.
Delegate questions drew out how customers actually operate the product. Akshay said most Clumio users work through the API or Terraform rather than the console, noting that the average user logs into the console only once every 37 days and that Clumio’s Terraform module was long among the most downloaded in data protection. Asked about recovery as code, he said Clumio has long described itself as data protection as code, and he noted that customers are now moving toward agentic recovery, in which an LLM orchestrates the entire restore from a prompt, though many still use scripted recovery. He also explained that Clumio describes its focus as AI-based rather than AI-native applications because it operates at the data layer, not at the agent or application level.
Jeff then gave a live demo in the Clumio SaaS interface. He showed a Google Cloud project he had already onboarded, a process he said takes about five minutes, and a bucket protected by a backup policy and a protection group. The policy sets how often backups run and how long they are kept, and the protection group applies the policy to a specific asset. The policy used continuous backup, which checks for changes every 15 minutes, the most frequent option for S3 and GCS. Jeff’s Python script simulated a ransomware attack on a prefix of 100 objects, renaming each one and replacing its 32-byte contents with a fake ransom note. Jeff then restored the entire prefix from the previous day’s Secure Vault copy, choosing the latest version only to save cost. In about two minutes, the object names, sizes, and contents were back to normal. He confirmed that GCS currently supports Secure Vault only, with no S3 Backtrack equivalent yet. He added that in a real attack he would recover to a separate Google Cloud project, since the rest of the compromised project may also be affected. On monitoring, Akshay said Clumio provides built-in task alerts but does not natively integrate with security or observability tools such as CrowdStrike, Wiz, or ServiceNow. Some customers connect those tools through Clumio’s API, which he called the product’s “gateway to the world.”
Personnel: Akshay Joshi, Jeff Cassianna
Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.
We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.
Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!