Tech Field Day

The Independent IT Influencer Event

  • Home
    • The Futurum Group
    • FAQ
    • Staff
  • Sponsors
    • Sponsor List
      • 2026 Sponsors
      • 2025 Sponsors
      • 2024 Sponsors
      • 2023 Sponsors
      • 2022 Sponsors
    • Sponsor Tech Field Day
    • Best of Tech Field Day
    • Results and Metrics
    • Preparing Your Presentation
      • Complete Presentation Guide
      • A Classic Tech Field Day Agenda
      • Field Day Room Setup
      • Presenting to Engineers
  • Delegates
    • Delegate List
      • 2026 Delegates
      • 2025 Delegates
      • 2024 Delegates
      • 2023 Delegates
      • 2022 Delegates
    • Become a Field Day Delegate
    • What Delegates Should Know
  • Events
    • All Events
      • Upcoming
      • Past
    • Field Day
    • Field Day Extra
    • Field Day Exclusive
    • Field Day Experience
    • Field Day Live
    • Field Day Showcase
  • Topics
    • Tech Field Day
    • Cloud Field Day
    • Mobility Field Day
    • Networking Field Day
    • Security Field Day
    • Storage Field Day
  • News
    • Coverage
    • Event News
    • Podcast
  • When autocomplete results are available use up and down arrows to review and enter to go to the desired page. Touch device users, explore by touch or with swipe gestures.
You are here: Home / Videos / Improving Deduplication via Mathematics with Richard Lary

Improving Deduplication via Mathematics with Richard Lary



Storage Field Day 13

Richard Lary presented for X-IO at SFD13




This video is part of the appearance, “X-IO Technologies Presents at Storage Field Day 13“. It was recorded as part of Storage Field Day 13 at 10:30-12:30 on June 16, 2017.


Watch on YouTube
Watch on Vimeo

In his presentation at Storage Field Day 13, Richard Lary, Chief Scientist at X-IO Technologies, delves into the complexities and innovations in deduplication technology, emphasizing the significant role of mathematics in enhancing performance. Lary begins by highlighting the resource-intensive nature of deduplication, which involves substantial memory, CPU cycles, and disk accesses. He explains that deduplication typically relies on computationally intensive signature techniques to compare incoming data with existing data, necessitating a robust and persistent database of signatures. This database must handle petascale systems with billions of entries, survive power and controller failures, and maintain high write throughput despite the challenges posed by the random nature of hash-based signatures. Lary critiques the inefficiency of traditional caching methods in this context and underscores the need for a high-throughput, persistent mapping database to manage the dynamic nature of deduplication.

Lary then introduces the concept of using non-crypto signatures, specifically MetroHash, to improve deduplication performance. He explains that while traditional crypto hashes like SHA-1 and SHA-256 are secure, they are slow and CPU-bound, making them less suitable for high-performance storage systems. In contrast, non-crypto hashes like MetroHash are significantly faster and can handle the high throughput demands of modern storage systems. Lary also discusses the innovative use of a “bouquet filter,” a variation of the bloom filter, to efficiently manage the deduplication process. This approach involves using multiple small bloom filters to reduce the computational overhead and improve performance. Lary hints at a proprietary method developed by X-IO to further optimize deduplication by treating unique data differently, thereby reducing wasted resources and enhancing overall system efficiency. This method, still under patent consideration, promises to significantly improve deduplication performance while maintaining data integrity.

Personnel: Richard Lary

  • Bluesky
  • LinkedIn
  • Mastodon
  • RSS
  • Twitter
  • YouTube

Event Calendar

  • Mar 23-Mar 24 — Tech Field Day Extra at RSAC 2026
  • Apr 8-Apr 10 — Networking Field Day 40
  • Apr 13-Apr 15 — Tech Field Day Experience at Qlik Connect 2026
  • Apr 29-Apr 30 — Security Field Day 15
  • May 6-May 8 — Mobility Field Day 14
  • May 13-May 14 — AI Field Day 8
  • Jun 2-Jun 3 — Tech Field Day Extra at Cisco Live US 2026
  • Jun 10-Jun 11 — AI Infrastructure Field Day 5

Latest Coverage

  • When Regulators Can’t Agree, Your Data Infrastructure Has to Carry the Weight
  • AI Guesses, Math Proves: Forward Networks Brings Deterministic Truth to AI Infrastructure Governance
  • When Storage Stops Being a Location
  • Qlik Answers, SpaceX vs Amazon, & Practical Quantum | Tech Field Day News Rundown: March 11, 2026
  • Preparing for CloudFieldDay 25

Tech Field Day News

  • The Frontlines of Cybersecurity at Tech Field Day Extra at RSAC 2026
  • Cloud Strategy, The Future of Infrastructure, and Of Course AI at Cloud Field Day 25

Return to top of page

Copyright © 2026 · Genesis Framework · WordPress · Log in