Return the k most frequent values in O(n) with counting plus bucket sort, beating the O(n log n) sort and O(n log k) heap. Asked at Amazon and Meta.
The problem
Return the k most frequent values in nums, most frequent first. The answer is unique (no ties at the cut-off). Aim for better than O(n log n).
Examples
Example 1
Input
top_k_frequent([1, 1, 1, 2, 2, 3], 2)
Expected output
[1, 2]
Example 2
Input
top_k_frequent([1], 1)
Expected output
[1]
+ 2 hidden tests on Submit.
Edge cases to ask about
- k = number of distinct values
- Negative numbers
Hints
0/3How an interviewer scores this
0/9Your code runs in real CPython inside your browser — nothing is sent anywhere. The first run downloads the interpreter (about 6 MB, once). Your code is saved on this device as you type.
Complexity Lab
What does this cost as n grows?
Interviewers score the analysis as much as the code. Commit to an answer first — then check it, and measure your code against the optimal one at growing input sizes.
Pick both to reveal the answer.
Measure it
Runs the function on inputs of size 250 up to 16,000 and records the time and peak memory. Slow solutions stop early — a short curve is itself the answer.
From brute force to optimal
The progression an interviewer wants to hear, one step at a time.
| Approach | Time | Space | Idea |
|---|---|---|---|
| Sort by count | O(n log n) | O(n) | |
| Heap of size k | O(n log k) | O(n) | Counter.most_common(k) uses heapq.nlargest. |
| bestBucket sort by count | O(n) | O(n) | A count can only be 1..n, so index buckets by count. |
Walkthrough of the optimal approach (try it yourself first)
Count with Counter. Every count lies between 1 and n, so create buckets indexed by count and drop each value into buckets[count]. Walking buckets from the highest count down yields the most frequent values first — O(n) overall, no comparison sort.
Counter(nums).most_common(k) (a heap, O(n log k)) is the idiomatic answer to mention first.
Complexity: O(n) time, O(n) space. Counting is O(n); there are n + 1 buckets and every value lands in exactly one.
Reveal the reference solution
from collections import Counter def top_k_frequent(nums, k): counts = Counter(nums) buckets = [[] for _ in range(len(nums) + 1)] for value, c in counts.items(): buckets[c].append(value) out = [] for c in range(len(buckets) - 1, 0, -1): for value in buckets[c]: out.append(value) if len(out) == k: return out return out
The brute force, for comparison
from collections import Counter def top_k_frequent(nums, k): return [v for v, _ in Counter(nums).most_common(k)]
Follow-ups interviewers ask
- Top k frequent WORDS, ties broken alphabetically.
- Stream version.
Frequently asked interview questions
Core interview concepts, complexities, and follow-ups scored by hiring teams.
What is the time complexity of Top K Frequent Elements in Python?
The optimal solution runs in O(n) time and O(n) auxiliary space. Counting is O(n); there are n + 1 buckets and every value lands in exactly one.
What is the brute-force approach, and how do you optimise it?
Sort by count: O(n log n) time, O(n) space. Heap of size k: O(n log k) time, O(n) space. Counter.most_common(k) uses heapq.nlargest. Bucket sort by count: O(n) time, O(n) space. A count can only be 1..n, so index buckets by count.
What follow-up questions do interviewers ask about Top K Frequent Elements?
Top k frequent WORDS, ties broken alphabetically. Stream version.
