An Incremental Multi-View Entity Resolution Framework for Dynamic Academic User Profiles in Big Data Environments

Authors

Keywords:

Record Linkage, Locality-Sensitive Hashing, Hierarchical Navigable Small World Graphs, Approximate Nearest Neighbor Search, Multi-View Representation, Incremental Indexing

Abstract

Entity Resolution (ER) in streaming Big Data environments demands dynamic, low-latency indexing architectures capable of resolving high-dimensional user profiles without incurring the computational overhead of complete index reconstruction. Conventional tree-based incremental indexing systems are fundamentally constrained by their inability to operate over semantic vector spaces, while Locality-Sensitive Hashing (LSH) blocking paradigms exhibit pronounced sensitivity to hash collisions, generating noisy candidate blocks that degrade downstream matching quality. This article proposes a unified, real-time incremental entity resolution framework engineered for multi-view academic user profiles. The proposed architecture integrates three synergistic components: (i) an incremental LSH-SimHash blocking layer that projects multi-view profile representations into candidate buckets while preserving high-dimensional angular relationships within a compact binary signature space; (ii) independent Hierarchical Navigable Small World (HNSW) graphs maintained within individual LSH buckets, functioning as localized approximate nearest neighbor (ANN) search engines that eliminate global structural reconstruction upon profile arrivals; and (iii) an Adaptive Candidate Selection (ACS) strategy that prunes retrieved neighbor sets using a multi-criteria scoring function jointly evaluating vector cosine similarity, graph neighborhood structural consistency, and cross-view semantic alignment. Comprehensive empirical evaluations conducted across three large-scale benchmark datasets—the Academic Profiles Dataset, the North Carolina Voter Registration Dataset, and the OZ Synthetic Dataset—demonstrate that the proposed framework achieves a peak F1-score of 96.84% while maintaining sub-4 ms end-to-end query resolution latencies and a stable, predictable memory footprint, consistently outperforming established tree-based and dynamic approximate blocking baselines across all evaluation dimensions.

Author Biographies

Atika Badaoui, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria

Doctoral student in Department of Computer Science, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria (2015); a Master’s degree in  in Fundamental Computer Science – University M'Hamed Bougara of Boumerdes, Algeria (2006).

Djamel Eddine Zegour, National School of Computer Science, Oued Smar, Algiers, Algeria,National School of Computer Science, Oued Smar, Algiers, Algeria

Professor at LCSI Laboratory (2010), National School of Computer Science ISI, Oued Smar, Algiers, Algeria (1989-2025). Postgraduation from Paris Dauphine University & INRIA (National Institute of Research in Computer Science and Automation), Paris, France (1988); a state doctorate, Algiers, Algeria (1998).

Walid Khaled Hidouci, National School of Computer Science ISI, Oued Smar, Algiers, Algeria

Professor at LCSI Laboratory (2010), National School of Computer Science ISI, Oued Smar, Algiers, Algeria. Master's Degree (Magister) from National School of Computer Science, Algiers, Algeria(1993). a state doctorate from National School of Computer Science, Algiers, Algeria (1998).

Amin Riad Maouche, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria

Professor at Department of Computer Science, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria (1999); Faculty of Electronics and Computer Science, University of Science and Technology Houari Boumediene, Algiers, Algeria (1999). PhD in Robotics from the Faculty of Electronics and Computer Science at the University of Science and Technology in Algiers, Algeria (2010).

References

Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., & Stefanidis, K. (2020). An overview of end-to-end entity resolution for big data. ACM Computing Surveys, 53(6), 1–42. https://doi.org/10.1145/3418896

Papadakis, G., Ioannou, E., Thanos, E., & Palpanas, T. (2021). The four generations of entity resolution. Synthesis Lectures on Data Management, 16(2), 1–170. https://doi.org/10.2200/S01150ED1V01Y202111DTM073

Elmagarmid, A. K., Ipeirotis, P. G., & Verykios, V. S. (2007). Duplicate record detection: A survey. IEEE Transactions on Knowledge and Data Engineering, 19(1), 1–16. https://doi.org/10.1109/TKDE.2007.250581

Papadakis, G., Skoutas, D., Thanos, E., & Palpanas, T. (2020). Blocking and filtering techniques for entity resolution: A survey. ACM Computing Surveys, 53(2), 1–42. https://doi.org/10.1145/3377455

Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer. https://doi.org/10.1007/978-3-642-31164-2

Kejriwal, M., & Miranker, D. P. (2015). An unsupervised algorithm for learning blocking schemes. In Proceedings of the 2015 IEEE International Conference on Data Mining (ICDM) (pp. 191–200). IEEE. https://doi.org/10.1109/ICDM.2015.58

Saeedi, A., Peukert, E., & Rahm, E. (2017). Comparative evaluation of distributed clustering schemes for multi-source entity resolution. In Proceedings of the 20th International Conference on Extending Database Technology (EDBT) (pp. 278–289). OpenProceedings. https://doi.org/10.5441/002/edbt.2017.26

Ramadan, B., Christen, P., Liang, H., & Ng, C. K. (2015). Dynamic sorted neighborhood indexing for real-time entity resolution. Journal of Data and Information Quality, 7(1–2), 1–28. https://doi.org/10.1145/2816821

Zhu, J., Wen, J., & Wu, J. (2019). Real-time entity resolution by multiple indices. In Proceedings of the 2019 IEEE International Conference on Big Data (pp. 1168–1173). IEEE. https://doi.org/10.1109/BigData47090.2019.9005984

Liang, H., Wang, Y., Christen, P., & Gayler, R. (2014, May). Noise-tolerant approximate blocking for dynamic real-time entity resolution. In Proceedings of the 18th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) (pp. 449–461). Springer. https://doi.org/10.1007/978-3-319-06605-9_37

Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghuveer, V. (2018). Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 ACM SIGMOD International Conference on Management of Data (pp. 19–34). ACM. https://doi.org/10.1145/3183713.3196926

Datar, M., Immorlica, N., Indyk, P., & Mirrokni, V. S. (2004). Locality-sensitive hashing scheme based on p-stable distributions. In Proceedings of the 20th Annual Symposium on Computational Geometry (SCG) (pp. 253–262). ACM. https://doi.org/10.1145/997817.997857

Winkler, W. E. (2006). Overview of Record Linkage and Current Research Directions (Research Report RRS2006/02). U.S. Census Bureau Statistical Research Division.

Li, Y., Li, J., Suhara, Y., Doan, A., & Tan, W. C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1), 50–60. https://doi.org/10.14778/3421424.3421431

Köpcke, H., Thor, A., & Rahm, E. (2010). Evaluation of entity resolution approaches on real-world match problems. Proceedings of the VLDB Endowment, 3(1–2), 484–493. https://doi.org/10.14778/1920841.1920904

Cohen, W., Ravikumar, P., & Fienberg, S. (2003). A comparison of string distance metrics for name-matching tasks. In Proceedings of the IJCAI-2003 Workshop on Information Integration on the Web (pp. 73–78). AAAI Press.

Tang, J., Zhang, J., Yao, L., Li, J., Zhang, L., & Su, Z. (2008). ArnetMiner: Extraction and mining of academic social networks. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 990–998). ACM. https://doi.org/10.1145/1401890.1402008

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410

Ferreira, A. A., Gonçalves, M. A., & Laender, A. H. F. (2012). A brief survey of automatic methods for author name disambiguation. ACM SIGMOD Record, 41(2), 15–26. https://doi.org/10.1145/2350036.2350040

Aggarwal, C. C., Hinneburg, A., & Keim, D. A. (2001). On the surprising behavior of distance metrics in high dimensional space. In Proceedings of the 8th International Conference on Database Theory (ICDT) (pp. 420–434). Springer. https://doi.org/10.1007/3-540-44503-X_27

Indyk, P., & Motwani, R. (1998). Approximate nearest neighbors: Towards removing the curse of dimensionality. In Proceedings of the 30th Annual ACM Symposium on Theory of Computing (STOC) (pp. 604–613). ACM. https://doi.org/10.1145/276698.276876

Charikar, M. S. (2002). Similarity estimation techniques from rounding algorithms. In Proceedings of the 34th Annual ACM Symposium on Theory of Computing (STOC) (pp. 380–388). ACM. https://doi.org/10.1145/509907.509965

Malkov, Y. A., & Yashunin, D. A. (2020). Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(4), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473

Bernhardsson, E. (2018). Annoy: Approximate Nearest Neighbors in C++/Python. GitHub. https://github.com/spotify/annoy

Yashunin Manku, G. S., Jain, A., & Das Sarma, A. (2007). Detecting near-duplicates for web crawling. In Proceedings of the 16th International Conference on World Wide Web (WWW) (pp. 141–150). ACM. https://doi.org/10.1145/1242572.1242592

Dong, W., Moses, C., & Li, K. (2011). Efficient k-nearest neighbor graph construction for generic similarity measures. In Proceedings of the 20th International Conference on World Wide Web (WWW) (pp. 577–586). ACM. https://doi.org/10.1145/1963405.1963487

Shen, W., Wang, J., Luo, P., & Wang, M. (2012). LINDEN: Linking named entities with knowledge base via semantic knowledge. In Proceedings of the 21st International Conference on World Wide Web (WWW) (pp. 449–458). ACM. https://doi.org/10.1145/2187836.2187898

Christen, P., & Pudjijono, A. (2009). Accurate synthetic generation of realistic personal information. In Proceedings of the 13th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) (pp. 507–514). Springer. https://doi.org/10.1007/978-3-642-01307-2_47

Downloads

Published

2026-08-04

How to Cite

Badaoui, A., Zegour, D. E., Hidouci, W. K., & Maouche, A. R. (2026). An Incremental Multi-View Entity Resolution Framework for Dynamic Academic User Profiles in Big Data Environments. Revista Coleta Científica, 10(20), e20267 . Retrieved from http://portalcoleta.com.br/index.php/rcc/article/view/267

Issue

Section

Management, Innovation and Technology

ARK