Abstract
This paper addresses the challenge of identifying super spreaders within large, high-speed data streams. In these streams, data is segmented into flows, with each flow’s spread defined as the number of distinct items it contains. A super spreader is characterized as a flow with a notably large spread. Measuring flow spread requires counting any item at its first appearance and ignoring its subsequent duplicate appearances. Current compact solutions, known as sketches, are designed to fit within the constrained memory of online devices. However, existing sketches face accuracy challenges in spread tracking due to the substantial memory needed to remove the impact of duplicate appearances for measuring a single flow’s spread—a problem that compounds with increasing flow counts. We propose a novel sketch-based solution to address these limitations. At its core, our approach features an innovative non-duplicate sampler that eliminates duplicate appearances of any item, enabling accurate flow spread calculation using simple counters. Combined with our exponential-weakening decay mechanism that emphasizes large flows, the solution significantly improves super spreader detection accuracy. We provide rigorous theoretical analysis of our method and validate its performance through trace-driven experiments. Results demonstrate that our approach statistically outperforms existing state-of-the-art solutions in super spreader identification. Moreover, it achieves the fastest super spreader restoration time and reduces bandwidth consumption by an order of magnitude during remote offline restoration.
| Original language | English |
|---|---|
| Pages (from-to) | 1484-1497 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Knowledge and Data Engineering |
| Volume | 38 |
| Issue number | 3 |
| DOIs | |
| State | Published - 2026 |
Bibliographical note
Publisher Copyright:© 1989-2012 IEEE.
Funding
This work was supported by the University of Kentucky Start-Up Fund.
| Funders |
|---|
| University of Kentucky |
Keywords
- Data streams
- non-duplicate sampling
- super spreaders identification
ASJC Scopus subject areas
- Information Systems
- Computer Science Applications
- Computational Theory and Mathematics
Fingerprint
Dive into the research topics of 'Accurate Super Spreader Identification With Non-Duplicate Samplers in Data Streams'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver