Computer Science > Data Structures and Algorithms
[Submitted on 10 Apr 2018 (v1), last revised 25 Sep 2019 (this version, v3)]
Title:Optimal Document Exchange and New Codes for Insertions and Deletions
View PDFAbstract:We give the first communication-optimal document exchange protocol. For any $n$ and $k < n$ our randomized scheme takes any $n$-bit file $F$ and computes a $\Theta(k \log \frac{n}{k})$-bit summary from which one can reconstruct $F$, with high probability, given a related file $F'$ with edit distance $ED(F,F') \leq k$.
The size of our summary is information-theoretically order optimal for all values of $k$, giving a randomized solution to a longstanding open question of [Orlitsky; FOCS'91]. It also is the first non-trivial solution for the interesting setting where a small constant fraction of symbols have been edited, producing an optimal summary of size $O(H(\delta)n)$ for $k=\delta n$. This concludes a long series of better-and-better protocols which produce larger summaries for sub-linear values of $k$ and sub-polynomial failure probabilities. In particular, the recent break-through of [Belazzougui, Zhang; FOCS'16] assumes that $k < n^\epsilon$, produces a summary of size $O(k\log^2 k + k\log n)$, and succeeds with probability $1-(k \log n)^{-O(1)}$.
We also give an efficient derandomized document exchange protocol with summary size $O(k \log^2 \frac{n}{k})$. This improves, for any $k$, over a deterministic document exchange protocol by Belazzougui with summary size $O(k^2 + k \log^2 n)$. Our deterministic document exchange directly provides new efficient systematic error correcting codes for insertions and deletions. These (binary) codes correct any $\delta$ fraction of adversarial insertions/deletions while having a rate of $1 - O(\delta \log^2 \frac{1}{\delta})$ and improve over the codes of Guruswami and Li and Haeupler, Shahrasbi and Vitercik which have rate $1 - \Theta\left(\sqrt{\delta} \log^{O(1)} \frac{1}{\epsilon}\right)$.
Submission history
From: Bernhard Haeupler [view email][v1] Tue, 10 Apr 2018 15:54:34 UTC (24 KB)
[v2] Mon, 17 Dec 2018 23:50:43 UTC (33 KB)
[v3] Wed, 25 Sep 2019 23:19:16 UTC (92 KB)
References & Citations
Bibliographic and Citation Tools
Bibliographic Explorer (What is the Explorer?)
Litmaps (What is Litmaps?)
scite Smart Citations (What are Smart Citations?)
Code, Data and Media Associated with this Article
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub (What is DagsHub?)
Gotit.pub (What is GotitPub?)
Papers with Code (What is Papers with Code?)
ScienceCast (What is ScienceCast?)
Demos
Recommenders and Search Tools
Influence Flower (What are Influence Flowers?)
Connected Papers (What is Connected Papers?)
CORE Recommender (What is CORE?)
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.