The Four Generations of Entity Resolution (Record no. 86140)
[ view plain ]
000 -LEADER | |
---|---|
fixed length control field | 04478nam a22005415i 4500 |
001 - CONTROL NUMBER | |
control field | 978-3-031-01878-7 |
005 - DATE AND TIME OF LATEST TRANSACTION | |
control field | 20240730165153.0 |
008 - FIXED-LENGTH DATA ELEMENTS--GENERAL INFORMATION | |
fixed length control field | 220601s2021 sz | s |||| 0|eng d |
020 ## - INTERNATIONAL STANDARD BOOK NUMBER | |
ISBN | 9783031018787 |
-- | 978-3-031-01878-7 |
082 04 - CLASSIFICATION NUMBER | |
Call Number | 004.6 |
100 1# - AUTHOR NAME | |
Author | Papadakis, George. |
245 14 - TITLE STATEMENT | |
Title | The Four Generations of Entity Resolution |
250 ## - EDITION STATEMENT | |
Edition statement | 1st ed. 2021. |
300 ## - PHYSICAL DESCRIPTION | |
Number of Pages | XVII, 152 p. |
490 1# - SERIES STATEMENT | |
Series statement | Synthesis Lectures on Data Management, |
505 0# - FORMATTED CONTENTS NOTE | |
Remark 2 | Preface -- Acknowledgments -- Entity Resolution: Past, Present, and Yet-to-Come -- Preliminaries -- Generation 1: Addressing Veracity -- Generation 2: Also Addressing Volume -- Generation 3: Also Addressing Variety -- Generation 4: Also Addressing Velocity -- Leveraging External Knowledge -- Resources for Entity Resolution -- Possible Directions for Future Work -- Bibliography -- Authors' Biographies. |
520 ## - SUMMARY, ETC. | |
Summary, etc | Entity Resolution (ER) lies at the core of data integration and cleaning and, thus, a bulk of the research examines ways for improving its effectiveness and time efficiency. The initial ER methods primarily target Veracity in the context of structured (relational) data that are described by a schema of well-known quality and meaning. To achieve high effectiveness, they leverage schema, expert, and/or external knowledge. Part of these methods are extended to address Volume, processing large datasets through multi-core or massive parallelization approaches, such as the MapReduce paradigm. However, these early schema-based approaches are inapplicable to Web Data, which abound in voluminous, noisy, semi-structured, and highly heterogeneous information. To address the additional challenge of Variety, recent works on ER adopt a novel, loosely schema-aware functionality that emphasizes scalability and robustness to noise. Another line of present research focuses on the additional challenge ofVelocity, aiming to process data collections of a continuously increasing volume. The latest works, though, take advantage of the significant breakthroughs in Deep Learning and Crowdsourcing, incorporating external knowledge to enhance the existing words to a significant extent. This synthesis lecture organizes ER methods into four generations based on the challenges posed by these four Vs. For each generation, we outline the corresponding ER workflow, discuss the state-of-the-art methods per workflow step, and present current research directions. The discussion of these methods takes into account a historical perspective, explaining the evolution of the methods over time along with their similarities and differences. The lecture also discusses the available ER tools and benchmark datasets that allow expert as well as novice users to make use of the available solutions. |
700 1# - AUTHOR 2 | |
Author 2 | Ioannou, Ekaterini. |
700 1# - AUTHOR 2 | |
Author 2 | Thanos, Emanouil. |
700 1# - AUTHOR 2 | |
Author 2 | Palpanas, Themis. |
856 40 - ELECTRONIC LOCATION AND ACCESS | |
Uniform Resource Identifier | https://doi.org/10.1007/978-3-031-01878-7 |
942 ## - ADDED ENTRY ELEMENTS (KOHA) | |
Koha item type | eBooks |
264 #1 - | |
-- | Cham : |
-- | Springer International Publishing : |
-- | Imprint: Springer, |
-- | 2021. |
336 ## - | |
-- | text |
-- | txt |
-- | rdacontent |
337 ## - | |
-- | computer |
-- | c |
-- | rdamedia |
338 ## - | |
-- | online resource |
-- | cr |
-- | rdacarrier |
347 ## - | |
-- | text file |
-- | |
-- | rda |
650 #0 - SUBJECT ADDED ENTRY--SUBJECT 1 | |
-- | Computer networks . |
650 #0 - SUBJECT ADDED ENTRY--SUBJECT 1 | |
-- | Data structures (Computer science). |
650 #0 - SUBJECT ADDED ENTRY--SUBJECT 1 | |
-- | Information theory. |
650 14 - SUBJECT ADDED ENTRY--SUBJECT 1 | |
-- | Computer Communication Networks. |
650 24 - SUBJECT ADDED ENTRY--SUBJECT 1 | |
-- | Data Structures and Information Theory. |
830 #0 - SERIES ADDED ENTRY--UNIFORM TITLE | |
-- | 2153-5426 |
912 ## - | |
-- | ZDB-2-SXSC |
No items available.