DSA Tracker

Hard

KMP Algorithm (Pattern Matching)

A hard Strings problem included in Love Babbar 450, Striver A2Z. Below: the roles whose interviews prioritise this topic, and how to practise it.

Topic
Strings
Sheets
2
Core for
13 roles
Platform
LeetCode

The problem

Implement the Knuth-Morris-Pratt algorithm to find all starting indices where the pattern occurs in the text.

Example 1

Input
haystack="ababcabcabababd", pattern="ababd"
Output
[10]

Example 2

Input
haystack="hello", pattern="ll"
Output
[2]

Example 3

Input
haystack="aaaaa", pattern="bba"
Output
[]

Constraints

  • 1 <= haystack.length, pattern.length <= 10^5
  • strings consist of lowercase English letters

How to think about it

Updated 2026-09-09

When finding every occurrence of a pattern, resetting after a complete match is just as wasteful as resetting on a mismatch. The prefix table tells you not only where to resume after a broken comparison, but also how much of the completed match can immediately serve as the beginning of an overlapping next occurrence.

Approaches, worst first

  1. Naive pattern scan

    time O(n * m) · space O(1)

    Test pattern alignment starting at each index in haystack from 0 to n - m. On repetitive strings like 'aaaaa' matching against 'aaa', inner comparisons are completely repeated, leading to quadratic worst-case runtimes.

  2. KMP with LPS tableWrite this one

    time O(n + m) · space O(m)

    Compute the longest prefix-suffix table for pattern in O(m). Stream through haystack: advance pattern pointer j on matching characters. When j reaches m, record the starting match index and shift j to lps[m - 1] to catch overlapping matches without rewinding the text pointer.

Where people lose marks · 3
  • Resetting the pattern pointer to 0 after finding a full match: this misses overlapping matches such as pattern 'aba' in text 'ababa'. Shift j to lps[m - 1] instead.
  • When computing the LPS table, forgetting to fall back to lps[prev - 1] iteratively when pattern[i] != pattern[prev] until a match is found or prev reaches 0.
  • Empty result array: when no occurrences exist, return an empty array rather than [-1] or null.

The theory behind it

Strings — the ground this problem stands on. All Strings problems

What Strings is

A string is an ordered necklace of text characters, like letters printed along a ribbon of paper. Each character sits at an exact numeric slot, holding a glyph such as a letter, punctuation mark, or digit. In many programming languages, ribbons cannot be edited after creation, meaning changing a single character requires pressing an entirely new ribbon from scratch.

When to reach for it

Reach for string techniques when inputs consist of words, DNA sequences, serialized data formats, or sentences. Clues include questions testing palindromes, anagram matches, substring patterns, parenthesis balancing, or character frequency counts. Whenever an algorithm asks to transform capitalization, parse structured tokens, or compute edits between two phrases, string representations are the core subject.

How the pattern works

Think of characters as small integer codes ranging across standard character sets. Frequency tables with fixed sizes often replace heavy hash maps when tallying occurrences. For search tasks, maintain rolling state using character indices or sliding borders. When building output text through repeated appends, accumulate pieces inside a mutable list or string builder rather than concatenating strings directly, avoiding quadratic copy overhead.

What each operation costs

OperationTime
read character by indexO(1)
concatenate two strings of total length nO(n)
compare two strings of length nO(n)
What usually goes wrong with Strings
  • Concatenating strings inside a loop using the plus operator, which silently creates full copies on each iteration and turns linear routines into quadratic slowdowns.
  • Assuming all characters fall strictly within lowercase English letters without validating spaces, uppercase variants, punctuation marks, or multi-byte unicode symbols.
  • Confusing substring length with end index when slicing, causing unexpected off-by-one truncations in languages that take length versus exclusive end position.

Which roles need this problem

Strings is a core topic for these 13 roles — if you're targeting one of them, this problem is early in your path, not optional.

Secondary for 7 more roles, including Data Engineer, Data Analyst, Embedded / Firmware Engineer.

Track this in your role's order

Pick your target role and all 370 problems — including this one — resequence to what that interview actually asks. Free.

Start free

More Strings problems

Problem set and role mapping as of .