Um Arcabouço Incremental de Resolução de Entidades Multivisão para Perfis Dinâmicos de Usuários Acadêmicos em Ambientes de Big Data

Autores

Palavras-chave:

Vinculação de Registros (Record Linkage), Hashing Sensível à Localidade, Grafos Hierarchical, Navigable Small World, Busca de Vizinhos Mais Próximos Aproximados, Representação Multivisão, Indexação Incremental

Resumo

A Resolução de Entidades (RE) em ambientes de Big Data em fluxo contínuo (streaming) exige arquiteturas de indexação dinâmicas e de baixa latência, capazes de resolver perfis de usuários de alta dimensão sem incorrer na sobrecarga computacional da reconstrução completa de índices. Sistemas convencionais de indexação incremental baseados em árvores são fundamentalmente limitados por sua incapacidade de operar sobre espaços vetoriais semânticos, enquanto os paradigmas de bloqueio por Hashing Sensível à Localidade (LSH) exibem uma acentuada sensibilidade a colisões de hash, gerando blocos candidatos ruidosos que degradam a qualidade do pareamento subsequente. Este artigo propõe um framework unificado e em tempo real de resolução incremental de entidades, projetado para perfis de usuários acadêmicos multivisão. A arquitetura proposta integra três componentes sinérgicos: (i) uma camada de bloqueio incremental LSH-SimHash que projeta representações de perfis multivisão em buckets candidatos, preservando relações angulares de alta dimensão em um espaço compacto de assinaturas binárias; (ii) grafos independentes Hierarchical Navigable Small World (HNSW) mantidos dentro de buckets LSH individuais, funcionando como motores localizados de busca por vizinhos mais próximos aproximados (ANN) que eliminam a reconstrução estrutural global a cada chegada de novos perfis; e (iii) uma estratégia de Seleção Adaptativa de Candidatos (Adaptive Candidate Selection - ACS) que poda os conjuntos de vizinhos recuperados utilizando uma função de pontuação multicritério que avalia conjuntamente a similaridade de cosseno vetorial, a consistência estrutural da vizinhança no grafo e o alinhamento semântico entre visões. Avaliações empíricas abrangentes realizadas em três conjuntos de dados (datasets) de referência em grande escala — o Academic Profiles Dataset, o North Carolina Voter Registration Dataset e o OZ Synthetic Dataset — demonstram que o framework proposto atinge um F1-score de pico de 96,84%, enquanto mantém latências de resolução de consultas ponta a ponta inferiores a 4 ms e um consumo de memória estável e previsível, superando consistentemente as linhas de base (baselines) estabelecidas baseadas em árvores e em bloqueio aproximado dinâmico em todas as dimensões de avaliação.

Biografia do Autor

Atika Badaoui, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria

Doctoral student in Department of Computer Science, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria (2015); a Master’s degree in  in Fundamental Computer Science – University M'Hamed Bougara of Boumerdes, Algeria (2006).

Djamel Eddine Zegour, National School of Computer Science, Oued Smar, Algiers, Algeria

Professor at LCSI Laboratory (2010), National School of Computer Science ISI, Oued Smar, Algiers, Algeria (1989-2025). Postgraduation from Paris Dauphine University & INRIA (National Institute of Research in Computer Science and Automation), Paris, France (1988); a state doctorate, Algiers, Algeria (1998).

Walid Khaled Hidouci, National School of Computer Science ISI, Oued Smar, Algiers, Algeria

Professor at LCSI Laboratory (2010), National School of Computer Science ISI, Oued Smar, Algiers, Algeria. Master's Degree (Magister) from National School of Computer Science, Algiers, Algeria(1993). a state doctorate from National School of Computer Science, Algiers, Algeria (1998).

Amin Riad Maouche, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria

Professor at Department of Computer Science, University M'Hamed Bougara of Boumerdes, Boumerdes, Algeria (1999); Faculty of Electronics and Computer Science, University of Science and Technology Houari Boumediene, Algiers, Algeria (1999). PhD in Robotics from the Faculty of Electronics and Computer Science at the University of Science and Technology in Algiers, Algeria (2010).

Referências

Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., & Stefanidis, K. (2020). An overview of end-to-end entity resolution for big data. ACM Computing Surveys, 53(6), 1–42. https://doi.org/10.1145/3418896

Papadakis, G., Ioannou, E., Thanos, E., & Palpanas, T. (2021). The four generations of entity resolution. Synthesis Lectures on Data Management, 16(2), 1–170. https://doi.org/10.2200/S01150ED1V01Y202111DTM073

Elmagarmid, A. K., Ipeirotis, P. G., & Verykios, V. S. (2007). Duplicate record detection: A survey. IEEE Transactions on Knowledge and Data Engineering, 19(1), 1–16. https://doi.org/10.1109/TKDE.2007.250581

Papadakis, G., Skoutas, D., Thanos, E., & Palpanas, T. (2020). Blocking and filtering techniques for entity resolution: A survey. ACM Computing Surveys, 53(2), 1–42. https://doi.org/10.1145/3377455

Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer. https://doi.org/10.1007/978-3-642-31164-2

Kejriwal, M., & Miranker, D. P. (2015). An unsupervised algorithm for learning blocking schemes. In Proceedings of the 2015 IEEE International Conference on Data Mining (ICDM) (pp. 191–200). IEEE. https://doi.org/10.1109/ICDM.2015.58

Saeedi, A., Peukert, E., & Rahm, E. (2017). Comparative evaluation of distributed clustering schemes for multi-source entity resolution. In Proceedings of the 20th International Conference on Extending Database Technology (EDBT) (pp. 278–289). OpenProceedings. https://doi.org/10.5441/002/edbt.2017.26

Ramadan, B., Christen, P., Liang, H., & Ng, C. K. (2015). Dynamic sorted neighborhood indexing for real-time entity resolution. Journal of Data and Information Quality, 7(1–2), 1–28. https://doi.org/10.1145/2816821

Zhu, J., Wen, J., & Wu, J. (2019). Real-time entity resolution by multiple indices. In Proceedings of the 2019 IEEE International Conference on Big Data (pp. 1168–1173). IEEE. https://doi.org/10.1109/BigData47090.2019.9005984

Liang, H., Wang, Y., Christen, P., & Gayler, R. (2014, May). Noise-tolerant approximate blocking for dynamic real-time entity resolution. In Proceedings of the 18th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) (pp. 449–461). Springer. https://doi.org/10.1007/978-3-319-06605-9_37

Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghuveer, V. (2018). Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 ACM SIGMOD International Conference on Management of Data (pp. 19–34). ACM. https://doi.org/10.1145/3183713.3196926

Datar, M., Immorlica, N., Indyk, P., & Mirrokni, V. S. (2004). Locality-sensitive hashing scheme based on p-stable distributions. In Proceedings of the 20th Annual Symposium on Computational Geometry (SCG) (pp. 253–262). ACM. https://doi.org/10.1145/997817.997857

Winkler, W. E. (2006). Overview of Record Linkage and Current Research Directions (Research Report RRS2006/02). U.S. Census Bureau Statistical Research Division.

Li, Y., Li, J., Suhara, Y., Doan, A., & Tan, W. C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1), 50–60. https://doi.org/10.14778/3421424.3421431

Köpcke, H., Thor, A., & Rahm, E. (2010). Evaluation of entity resolution approaches on real-world match problems. Proceedings of the VLDB Endowment, 3(1–2), 484–493. https://doi.org/10.14778/1920841.1920904

Cohen, W., Ravikumar, P., & Fienberg, S. (2003). A comparison of string distance metrics for name-matching tasks. In Proceedings of the IJCAI-2003 Workshop on Information Integration on the Web (pp. 73–78). AAAI Press.

Tang, J., Zhang, J., Yao, L., Li, J., Zhang, L., & Su, Z. (2008). ArnetMiner: Extraction and mining of academic social networks. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 990–998). ACM. https://doi.org/10.1145/1401890.1402008

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410

Ferreira, A. A., Gonçalves, M. A., & Laender, A. H. F. (2012). A brief survey of automatic methods for author name disambiguation. ACM SIGMOD Record, 41(2), 15–26. https://doi.org/10.1145/2350036.2350040

Aggarwal, C. C., Hinneburg, A., & Keim, D. A. (2001). On the surprising behavior of distance metrics in high dimensional space. In Proceedings of the 8th International Conference on Database Theory (ICDT) (pp. 420–434). Springer. https://doi.org/10.1007/3-540-44503-X_27

Indyk, P., & Motwani, R. (1998). Approximate nearest neighbors: Towards removing the curse of dimensionality. In Proceedings of the 30th Annual ACM Symposium on Theory of Computing (STOC) (pp. 604–613). ACM. https://doi.org/10.1145/276698.276876

Charikar, M. S. (2002). Similarity estimation techniques from rounding algorithms. In Proceedings of the 34th Annual ACM Symposium on Theory of Computing (STOC) (pp. 380–388). ACM. https://doi.org/10.1145/509907.509965

Malkov, Y. A., & Yashunin, D. A. (2020). Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(4), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473

Bernhardsson, E. (2018). Annoy: Approximate Nearest Neighbors in C++/Python. GitHub. https://github.com/spotify/annoy

Yashunin Manku, G. S., Jain, A., & Das Sarma, A. (2007). Detecting near-duplicates for web crawling. In Proceedings of the 16th International Conference on World Wide Web (WWW) (pp. 141–150). ACM. https://doi.org/10.1145/1242572.1242592

Dong, W., Moses, C., & Li, K. (2011). Efficient k-nearest neighbor graph construction for generic similarity measures. In Proceedings of the 20th International Conference on World Wide Web (WWW) (pp. 577–586). ACM. https://doi.org/10.1145/1963405.1963487

Shen, W., Wang, J., Luo, P., & Wang, M. (2012). LINDEN: Linking named entities with knowledge base via semantic knowledge. In Proceedings of the 21st International Conference on World Wide Web (WWW) (pp. 449–458). ACM. https://doi.org/10.1145/2187836.2187898

Christen, P., & Pudjijono, A. (2009). Accurate synthetic generation of realistic personal information. In Proceedings of the 13th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) (pp. 507–514). Springer. https://doi.org/10.1007/978-3-642-01307-2_47

Downloads

Publicado

2026-08-04

Como Citar

Badaoui, A., Zegour, D. E., Hidouci, W. K., & Maouche, A. R. (2026). Um Arcabouço Incremental de Resolução de Entidades Multivisão para Perfis Dinâmicos de Usuários Acadêmicos em Ambientes de Big Data. Revista Coleta Científica, 10(20), e20267 . Recuperado de http://portalcoleta.com.br/index.php/rcc/article/view/267

Edição

Seção

Gestão, Inovação e Tecnologia

ARK