article Open access

A Hybrid Length-Based Pattern Matching Algorithm for Text Searching

  • International Journal of Advanced Computer Science and Applications
  • Science and Information Organization
Research footprint

At a glance

Citations
0
References
8
Comments
0
Paper overview

Abstract

This paper presents a hybrid algorithm for pattern matching in text, which combines word length preprocessing with the Knuth-Morris-Pratt (KMP) algorithm. Its performance was evaluated against KMP and Boyer-Moore (BM) in two scenarios: synthetic texts and real-world texts. In the former, classical algorithms proved more efficient due to the uniform structure of the data. However, in real-world texts, the hybrid algorithm significantly reduced search times, thanks to its ability to filter matches by length patterns before performing character-by-character comparisons. The algorithm also demonstrated flexibility in recognizing patterns with different delimiters. Among its limitations is the difficulty in detecting substrings within longer words. As future work, the incorporation of partial matching techniques and the adaptation of the approach to multilingual environments and machine learning systems are proposed. The dataset used is provided to encourage reproducibility.

Record transparency

Publication details

DOI
10.14569/ijacsa.2025.0160407
OpenAlex
W4410062632
Document type
article
Language
EN
Source
International Journal of Advanced Computer Science and Applications
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.