Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks (2206.12733v2)

Published 25 Jun 2022 in cs.DB

Abstract: How can we leverage existing column relationships within silos, to predict similar ones across silos? Can we do this efficiently and effectively? Existing matching approaches do not exploit prior knowledge, relying on prohibitively expensive similarity computations. In this paper we present the first technique for matching columns across data silos, called SiMa, which leverages Graph Neural Networks (GNNs) to learn from existing column relationships within data silos, and dataset-specific profiles. The main novelty of SiMa is its ability to be trained incrementally on column relationships within each silo individually, without requiring the consolidation of all datasets in a single place. Our experiments show that SiMa is more effective than the - otherwise inapplicable to the setting of silos - state-of-the-art matching methods, while requiring orders of magnitude less computational resources. Moreover, we demonstrate that SiMa considerably outperforms other state-of-the-art column representation learning methods.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Christos Koutras (5 papers)
  2. Rihan Hai (17 papers)
  3. Kyriakos Psarakis (7 papers)
  4. Marios Fragkoulis (9 papers)
  5. Asterios Katsifodimos (19 papers)
Citations (1)

Summary

We haven't generated a summary for this paper yet.