L2 · Working engineerAdvanced L2~8 min · 3 tests#90

Process a Huge File Without Loading It

Sum amounts from a huge CSV-style stream by iterating line by line instead of readlines(), keeping memory O(1). Verified with a memory-limit test.

The problem

lines is an iterator over the lines of a huge file (think 10 GB), each "id,amount". Return the total amount as an int.

You may only iterate once, and you must not hold all lines in memory: the test feeds 200,000 lines and fails if peak memory goes above 1 MB.

Examples

  1. Example 1

    Input

    total_amount(['1,10\n', '2,20\n', '3,30\n'])

    Expected output

    60
  2. Example 2

    Input

    check_streaming(total_amount, 200000)

    Expected output

    (599994, True)

+ 1 hidden test on Submit.

Edge cases to ask about

  • Empty file
  • Trailing newline
How the tests call your code

These helpers run before your code. The test inputs above call them.

import tracemalloc as _tm

def stream(n):
    for i in range(n):
        yield f"{i},{i % 7}\n"

def check_streaming(fn, n):
    _tm.start()
    try:
        total = fn(stream(n))
        peak = _tm.get_traced_memory()[1]
    finally:
        _tm.stop()
    return total, peak < 1024 * 1024

Hints

0/3

    How an interviewer scores this

    0/9
    Python 3.13 · total_amount
    ⌘/Ctrl + Enter runs the examples

    Your code runs in real CPython inside your browser — nothing is sent anywhere. The first run downloads the interpreter (about 6 MB, once). Your code is saved on this device as you type.

    Complexity Lab

    What does this cost as n grows?

    Interviewers score the analysis as much as the code. Commit to an answer first — then check it, and read why.

    Time complexity of the optimal solution
    Space complexity (extra memory)

    Pick both to reveal the answer.

    From brute force to optimal

    The progression an interviewer wants to hear, one step at a time.

    ApproachTimeSpaceIdea
    f.readlines() / list(lines)O(n)O(n)Loads the whole file — dies on 10 GB.
    bestIterate the file objectO(n)O(1)A file is a lazy iterator of lines; keep only a running total.
    Walkthrough of the optimal approach (try it yourself first)

    A Python file object is a lazy iterator: for line in f reads a buffered chunk at a time, so memory stays constant regardless of file size. f.read(), f.readlines() and list(f) all load everything.

    For binary or fixed-size processing, read in chunks: for chunk in iter(lambda: f.read(1 << 20), b""). For parallelism across a huge file, split by byte offsets and align to the next newline.

    Complexity: O(n) time, O(1) space. Each line is processed and discarded; only the running total stays in memory.

    Reveal the reference solution
    def total_amount(lines):
        total = 0
        for line in lines:
            _id, amount = line.rstrip("\n").split(",")
            total += int(amount)
        return total

    Follow-ups interviewers ask

    • Process it in parallel across 8 cores.
    • The file is gzip-compressed.

    Frequently asked interview questions

    Core interview concepts, complexities, and follow-ups scored by hiring teams.

    What is the time complexity of Process a Huge File Without Loading It in Python?

    The optimal solution runs in O(n) time and O(1) auxiliary space. Each line is processed and discarded; only the running total stays in memory.

    What is the brute-force approach, and how do you optimise it?

    f.readlines() / list(lines): O(n) time, O(n) space. Loads the whole file — dies on 10 GB. Iterate the file object: O(n) time, O(1) space. A file is a lazy iterator of lines; keep only a running total.

    What follow-up questions do interviewers ask about Process a Huge File Without Loading It?

    Process it in parallel across 8 cores. The file is gzip-compressed.