Watch on YouTube
Watch on Vimeo
Cloud-native applications depend on data distributed across databases, object stores, retrieval pipelines, and lakehouse tables. When that data is deleted, corrupted, or changed by a bad deployment, recovery can require full restores, temporary resources, application reconfiguration, and rebuilding downstream dependencies. In this Cloud Field Day 26 session, Jeff Cassianna and his colleague Akshay used a fictitious streaming app called Clumio Final Cut to show what that looks like for an AI feature. Its “AI Movie Assistant” concierge chatbot is a typical LLM application backed by a retrieval-augmented generation (RAG) pipeline with three layers: movie metadata stored as JSON, text, and PDF objects in an Amazon S3 bucket, vector embeddings of that data held in S3 Vectors, and the LLM that queries the vector store to answer questions. When the underlying S3 objects are deleted or corrupted, the embeddings point to nothing. The chatbot then either fails or starts hallucinating, for example, claiming that Love Actually is an action movie and failing to find the film “Clumio to the Moon.”
Akshay framed recovery around the SRE’s priorities. The team needs a recovery point with minimal data loss, as little LLM reconfiguration as possible (a slow and error-prone process), controlled costs, and minimal downtime. With traditional all-or-nothing tools, recovery takes five steps: restore the entire bucket to a new location, recompute the vectors, delete the old vector store, clean up the old bucket, and repoint the LLM. Clumio Backtrack for S3, launched at re:Invent 2024 after the Commvault acquisition, reduces this to selecting a point in time and the affected objects, then clicking restore. Backtrack combines three capabilities. It offers near-instant point-in-time rollback down to the second, granular recovery at the object, prefix, or bucket level, and in-place recovery, backed by continuous change tracking of hundreds of millions of object changes per hour. Responding to delegate questions, Akshay said in-place recovery suits internal incidents such as a bad code push rather than malicious attacks. He also said granular recovery effectively merges good data with restored data, and Clumio does not currently detect when an event occurred, so customers must bring that timestamp themselves. He argued that backup is now commoditized and that Clumio is investing in recovery options.
Jeff’s demo corrupted the “Clumio to the Moon” JSON object with a Python script, showed the chatbot returning inaccurate answers, and then used Backtrack to locate and overwrite the object with a prior version in under a minute, after which the chatbot returned accurate responses. Jeff noted that Backtrack moves no data; it makes S3 versioning usable at the scale of billions of objects. Secure Vault is Clumio’s isolated, air-gapped offering that copies data outside the customer’s security sphere. Backtrack requires versioning when used without Secure Vault. You can run bulk recoveries of thousands or billions of objects through the API, and Clumio has been tested with tens of billions of S3 objects and hundreds of billions of DynamoDB items. Its serverless, Lambda-based architecture is designed to avoid throttling. Restores can target a different onboarded AWS account or region, which Jeff recommended for cyber events or regional outages like the recent us-east-1 incident and which also helps with data sovereignty requirements. He closed by noting that Clumio recently achieved FedRAMP Ready status, with further compliance details at trust.commvault.com.
Personnel: Akshay Joshi, Jeff Cassianna
Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.
We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.
Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!