Chain Python generators into a lazy read → filter even → square pipeline that streams one item at a time in O(1) memory. Each stage is tested.
The problem
Write three generator functions and chain them:
read(source)— yields each item ofsourceonly_even(stream)— yields the even numberssquared(stream)— yields each number squared
and pipeline(source) that returns list(squared(only_even(read(source)))).
Examples
Example 1
Input
pipeline([1, 2, 3, 4, 5])
Expected output
[4, 16]
Example 2
Input
all_generators(read, only_even, squared)
Expected output
True
+ 2 hidden tests on Submit — pulls only what it needs.
Edge cases to ask about
- Empty source
- Laziness: stop early
How the tests call your code
These helpers run before your code. The test inputs above call them.
import inspect as _inspect def all_generators(read, only_even, squared): return all(_inspect.isgenerator(g) for g in (read([1]), only_even(iter([2])), squared(iter([2])))) def lazy_check(read, only_even, squared): pulled = [] def source(): for i in range(1, 1_000_000): pulled.append(i) yield i first = next(squared(only_even(read(source())))) return first, len(pulled)
Hints
0/3How an interviewer scores this
0/9Your code runs in real CPython inside your browser — nothing is sent anywhere. The first run downloads the interpreter (about 6 MB, once). Your code is saved on this device as you type.
Complexity Lab
What does this cost as n grows?
Interviewers score the analysis as much as the code. Commit to an answer first — then check it, and measure your code against the optimal one at growing input sizes.
Pick both to reveal the answer.
Measure it
Runs the function on inputs of size 250 up to 16,000 and records the time and peak memory. Slow solutions stop early — a short curve is itself the answer.
From brute force to optimal
The progression an interviewer wants to hear, one step at a time.
| Approach | Time | Space | Idea |
|---|---|---|---|
| Lists at each stage | O(n) | O(n) | Every stage materialises its whole output. |
| bestChained generators | O(n) | O(1) | Items flow through one at a time, on demand. |
Walkthrough of the optimal approach (try it yourself first)
Each stage is a generator that consumes the previous one. Nothing runs until something pulls from the end — the hidden test asks for just the first result and checks that only two source items were ever read.
That laziness is how you process files bigger than memory, infinite streams, or paginated APIs. Generator expressions make the same pipeline in one line: (x * x for x in source if x % 2 == 0).
Complexity: O(n) time, O(1) space. Each item passes through every stage once; no stage holds more than one item at a time (the final list() aside).
Reveal the reference solution
def read(source): for item in source: yield item def only_even(stream): for x in stream: if x % 2 == 0: yield x def squared(stream): for x in stream: yield x * x def pipeline(source): return list(squared(only_even(read(source))))
Follow-ups interviewers ask
- Add a batch(stream, size) stage.
- What happens if you iterate a generator twice?
Frequently asked interview questions
Core interview concepts, complexities, and follow-ups scored by hiring teams.
What is the time complexity of Build a Generator Data Pipeline in Python?
The optimal solution runs in O(n) time and O(1) auxiliary space. Each item passes through every stage once; no stage holds more than one item at a time (the final list() aside).
What is the brute-force approach, and how do you optimise it?
Lists at each stage: O(n) time, O(n) space. Every stage materialises its whole output. Chained generators: O(n) time, O(1) space. Items flow through one at a time, on demand.
What follow-up questions do interviewers ask about Build a Generator Data Pipeline?
Add a batch(stream, size) stage. What happens if you iterate a generator twice?
