KMP Algorithm (Pattern Matching)
A hard Strings problem included in Love Babbar 450, Striver A2Z. Below: the roles whose interviews prioritise this topic, and how to practise it.
- Topic
- Strings
- Sheets
- 2
- Core for
- 13 roles
- Platform
- LeetCode
The problem
Implement the Knuth-Morris-Pratt algorithm to find all starting indices where the pattern occurs in the text.
Example 1
- Input
- haystack="ababcabcabababd", pattern="ababd"
- Output
- [10]
Example 2
- Input
- haystack="hello", pattern="ll"
- Output
- [2]
Example 3
- Input
- haystack="aaaaa", pattern="bba"
- Output
- []
Constraints
- 1 <= haystack.length, pattern.length <= 10^5
- strings consist of lowercase English letters
How to think about it
Updated 2026-09-09When finding every occurrence of a pattern, resetting after a complete match is just as wasteful as resetting on a mismatch. The prefix table tells you not only where to resume after a broken comparison, but also how much of the completed match can immediately serve as the beginning of an overlapping next occurrence.
Approaches, worst first
Naive pattern scan
time O(n * m) · space O(1)
Test pattern alignment starting at each index in haystack from 0 to n - m. On repetitive strings like 'aaaaa' matching against 'aaa', inner comparisons are completely repeated, leading to quadratic worst-case runtimes.
KMP with LPS tableWrite this one
time O(n + m) · space O(m)
Compute the longest prefix-suffix table for pattern in O(m). Stream through haystack: advance pattern pointer j on matching characters. When j reaches m, record the starting match index and shift j to lps[m - 1] to catch overlapping matches without rewinding the text pointer.
Where people lose marks · 3
- Resetting the pattern pointer to 0 after finding a full match: this misses overlapping matches such as pattern 'aba' in text 'ababa'. Shift j to lps[m - 1] instead.
- When computing the LPS table, forgetting to fall back to lps[prev - 1] iteratively when pattern[i] != pattern[prev] until a match is found or prev reaches 0.
- Empty result array: when no occurrences exist, return an empty array rather than [-1] or null.
The theory behind it
Strings — the ground this problem stands on. All Strings problems
What Strings is
A string is an ordered necklace of text characters, like letters printed along a ribbon of paper. Each character sits at an exact numeric slot, holding a glyph such as a letter, punctuation mark, or digit. In many programming languages, ribbons cannot be edited after creation, meaning changing a single character requires pressing an entirely new ribbon from scratch.
When to reach for it
Reach for string techniques when inputs consist of words, DNA sequences, serialized data formats, or sentences. Clues include questions testing palindromes, anagram matches, substring patterns, parenthesis balancing, or character frequency counts. Whenever an algorithm asks to transform capitalization, parse structured tokens, or compute edits between two phrases, string representations are the core subject.
How the pattern works
Think of characters as small integer codes ranging across standard character sets. Frequency tables with fixed sizes often replace heavy hash maps when tallying occurrences. For search tasks, maintain rolling state using character indices or sliding borders. When building output text through repeated appends, accumulate pieces inside a mutable list or string builder rather than concatenating strings directly, avoiding quadratic copy overhead.
What each operation costs
| Operation | Time |
|---|---|
| read character by index | O(1) |
| concatenate two strings of total length n | O(n) |
| compare two strings of length n | O(n) |
What usually goes wrong with Strings
- Concatenating strings inside a loop using the plus operator, which silently creates full copies on each iteration and turns linear routines into quadratic slowdowns.
- Assuming all characters fall strictly within lowercase English letters without validating spaces, uppercase variants, punctuation marks, or multi-byte unicode symbols.
- Confusing substring length with end index when slicing, causing unexpected off-by-one truncations in languages that take length versus exclusive end position.
Which roles need this problem
Strings is a core topic for these 13 roles — if you're targeting one of them, this problem is early in your path, not optional.
Secondary for 7 more roles, including Data Engineer, Data Analyst, Embedded / Firmware Engineer.
Track this in your role's order
Pick your target role and all 370 problems — including this one — resequence to what that interview actually asks. Free.
Start freeMore Strings problems
Problem set and role mapping as of .