Skip to main navigation Skip to search Skip to main content

Accurate Super Spreader Identification With Non-Duplicate Samplers in Data Streams

Research output: Contribution to journalArticlepeer-review

Abstract

This paper addresses the challenge of identifying super spreaders within large, high-speed data streams. In these streams, data is segmented into flows, with each flow’s spread defined as the number of distinct items it contains. A super spreader is characterized as a flow with a notably large spread. Measuring flow spread requires counting any item at its first appearance and ignoring its subsequent duplicate appearances. Current compact solutions, known as sketches, are designed to fit within the constrained memory of online devices. However, existing sketches face accuracy challenges in spread tracking due to the substantial memory needed to remove the impact of duplicate appearances for measuring a single flow’s spread—a problem that compounds with increasing flow counts. We propose a novel sketch-based solution to address these limitations. At its core, our approach features an innovative non-duplicate sampler that eliminates duplicate appearances of any item, enabling accurate flow spread calculation using simple counters. Combined with our exponential-weakening decay mechanism that emphasizes large flows, the solution significantly improves super spreader detection accuracy. We provide rigorous theoretical analysis of our method and validate its performance through trace-driven experiments. Results demonstrate that our approach statistically outperforms existing state-of-the-art solutions in super spreader identification. Moreover, it achieves the fastest super spreader restoration time and reduces bandwidth consumption by an order of magnitude during remote offline restoration.

Original languageEnglish
Pages (from-to)1484-1497
Number of pages14
JournalIEEE Transactions on Knowledge and Data Engineering
Volume38
Issue number3
DOIs
StatePublished - 2026

Bibliographical note

Publisher Copyright:
© 1989-2012 IEEE.

Funding

This work was supported by the University of Kentucky Start-Up Fund.

Funders
University of Kentucky

    Keywords

    • Data streams
    • non-duplicate sampling
    • super spreaders identification

    ASJC Scopus subject areas

    • Information Systems
    • Computer Science Applications
    • Computational Theory and Mathematics

    Fingerprint

    Dive into the research topics of 'Accurate Super Spreader Identification With Non-Duplicate Samplers in Data Streams'. Together they form a unique fingerprint.

    Cite this