An Incremental Multi-View Entity Resolution Framework for Dynamic Academic User Profiles in Big Data Environments
Keywords:
Record Linkage, Locality-Sensitive Hashing, Hierarchical Navigable Small World Graphs, Approximate Nearest Neighbor Search, Multi-View Representation, Incremental IndexingAbstract
Entity Resolution (ER) in streaming Big Data environments demands dynamic, low-latency indexing architectures capable of resolving high-dimensional user profiles without incurring the computational overhead of complete index reconstruction. Conventional tree-based incremental indexing systems are fundamentally constrained by their inability to operate over semantic vector spaces, while Locality-Sensitive Hashing (LSH) blocking paradigms exhibit pronounced sensitivity to hash collisions, generating noisy candidate blocks that degrade downstream matching quality. This article proposes a unified, real-time incremental entity resolution framework engineered for multi-view academic user profiles. The proposed architecture integrates three synergistic components: (i) an incremental LSH-SimHash blocking layer that projects multi-view profile representations into candidate buckets while preserving high-dimensional angular relationships within a compact binary signature space; (ii) independent Hierarchical Navigable Small World (HNSW) graphs maintained within individual LSH buckets, functioning as localized approximate nearest neighbor (ANN) search engines that eliminate global structural reconstruction upon profile arrivals; and (iii) an Adaptive Candidate Selection (ACS) strategy that prunes retrieved neighbor sets using a multi-criteria scoring function jointly evaluating vector cosine similarity, graph neighborhood structural consistency, and cross-view semantic alignment. Comprehensive empirical evaluations conducted across three large-scale benchmark datasets—the Academic Profiles Dataset, the North Carolina Voter Registration Dataset, and the OZ Synthetic Dataset—demonstrate that the proposed framework achieves a peak F1-score of 96.84% while maintaining sub-4 ms end-to-end query resolution latencies and a stable, predictable memory footprint, consistently outperforming established tree-based and dynamic approximate blocking baselines across all evaluation dimensions.
References
Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., & Stefanidis, K. (2020). An overview of end-to-end entity resolution for big data. ACM Computing Surveys, 53(6), 1–42. https://doi.org/10.1145/3418896
Papadakis, G., Ioannou, E., Thanos, E., & Palpanas, T. (2021). The four generations of entity resolution. Synthesis Lectures on Data Management, 16(2), 1–170. https://doi.org/10.2200/S01150ED1V01Y202111DTM073
Elmagarmid, A. K., Ipeirotis, P. G., & Verykios, V. S. (2007). Duplicate record detection: A survey. IEEE Transactions on Knowledge and Data Engineering, 19(1), 1–16. https://doi.org/10.1109/TKDE.2007.250581
Papadakis, G., Skoutas, D., Thanos, E., & Palpanas, T. (2020). Blocking and filtering techniques for entity resolution: A survey. ACM Computing Surveys, 53(2), 1–42. https://doi.org/10.1145/3377455
Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer. https://doi.org/10.1007/978-3-642-31164-2
Kejriwal, M., & Miranker, D. P. (2015). An unsupervised algorithm for learning blocking schemes. In Proceedings of the 2015 IEEE International Conference on Data Mining (ICDM) (pp. 191–200). IEEE. https://doi.org/10.1109/ICDM.2015.58
Saeedi, A., Peukert, E., & Rahm, E. (2017). Comparative evaluation of distributed clustering schemes for multi-source entity resolution. In Proceedings of the 20th International Conference on Extending Database Technology (EDBT) (pp. 278–289). OpenProceedings. https://doi.org/10.5441/002/edbt.2017.26
Ramadan, B., Christen, P., Liang, H., & Ng, C. K. (2015). Dynamic sorted neighborhood indexing for real-time entity resolution. Journal of Data and Information Quality, 7(1–2), 1–28. https://doi.org/10.1145/2816821
Zhu, J., Wen, J., & Wu, J. (2019). Real-time entity resolution by multiple indices. In Proceedings of the 2019 IEEE International Conference on Big Data (pp. 1168–1173). IEEE. https://doi.org/10.1109/BigData47090.2019.9005984
Liang, H., Wang, Y., Christen, P., & Gayler, R. (2014, May). Noise-tolerant approximate blocking for dynamic real-time entity resolution. In Proceedings of the 18th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) (pp. 449–461). Springer. https://doi.org/10.1007/978-3-319-06605-9_37
Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghuveer, V. (2018). Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 ACM SIGMOD International Conference on Management of Data (pp. 19–34). ACM. https://doi.org/10.1145/3183713.3196926
Datar, M., Immorlica, N., Indyk, P., & Mirrokni, V. S. (2004). Locality-sensitive hashing scheme based on p-stable distributions. In Proceedings of the 20th Annual Symposium on Computational Geometry (SCG) (pp. 253–262). ACM. https://doi.org/10.1145/997817.997857
Winkler, W. E. (2006). Overview of Record Linkage and Current Research Directions (Research Report RRS2006/02). U.S. Census Bureau Statistical Research Division.
Li, Y., Li, J., Suhara, Y., Doan, A., & Tan, W. C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1), 50–60. https://doi.org/10.14778/3421424.3421431
Köpcke, H., Thor, A., & Rahm, E. (2010). Evaluation of entity resolution approaches on real-world match problems. Proceedings of the VLDB Endowment, 3(1–2), 484–493. https://doi.org/10.14778/1920841.1920904
Cohen, W., Ravikumar, P., & Fienberg, S. (2003). A comparison of string distance metrics for name-matching tasks. In Proceedings of the IJCAI-2003 Workshop on Information Integration on the Web (pp. 73–78). AAAI Press.
Tang, J., Zhang, J., Yao, L., Li, J., Zhang, L., & Su, Z. (2008). ArnetMiner: Extraction and mining of academic social networks. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 990–998). ACM. https://doi.org/10.1145/1401890.1402008
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410
Ferreira, A. A., Gonçalves, M. A., & Laender, A. H. F. (2012). A brief survey of automatic methods for author name disambiguation. ACM SIGMOD Record, 41(2), 15–26. https://doi.org/10.1145/2350036.2350040
Aggarwal, C. C., Hinneburg, A., & Keim, D. A. (2001). On the surprising behavior of distance metrics in high dimensional space. In Proceedings of the 8th International Conference on Database Theory (ICDT) (pp. 420–434). Springer. https://doi.org/10.1007/3-540-44503-X_27
Indyk, P., & Motwani, R. (1998). Approximate nearest neighbors: Towards removing the curse of dimensionality. In Proceedings of the 30th Annual ACM Symposium on Theory of Computing (STOC) (pp. 604–613). ACM. https://doi.org/10.1145/276698.276876
Charikar, M. S. (2002). Similarity estimation techniques from rounding algorithms. In Proceedings of the 34th Annual ACM Symposium on Theory of Computing (STOC) (pp. 380–388). ACM. https://doi.org/10.1145/509907.509965
Malkov, Y. A., & Yashunin, D. A. (2020). Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(4), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473
Bernhardsson, E. (2018). Annoy: Approximate Nearest Neighbors in C++/Python. GitHub. https://github.com/spotify/annoy
Yashunin Manku, G. S., Jain, A., & Das Sarma, A. (2007). Detecting near-duplicates for web crawling. In Proceedings of the 16th International Conference on World Wide Web (WWW) (pp. 141–150). ACM. https://doi.org/10.1145/1242572.1242592
Dong, W., Moses, C., & Li, K. (2011). Efficient k-nearest neighbor graph construction for generic similarity measures. In Proceedings of the 20th International Conference on World Wide Web (WWW) (pp. 577–586). ACM. https://doi.org/10.1145/1963405.1963487
Shen, W., Wang, J., Luo, P., & Wang, M. (2012). LINDEN: Linking named entities with knowledge base via semantic knowledge. In Proceedings of the 21st International Conference on World Wide Web (WWW) (pp. 449–458). ACM. https://doi.org/10.1145/2187836.2187898
Christen, P., & Pudjijono, A. (2009). Accurate synthetic generation of realistic personal information. In Proceedings of the 13th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) (pp. 507–514). Springer. https://doi.org/10.1007/978-3-642-01307-2_47
















