L2 · Working engineerAdvanced L2~6 min · 4 tests#91

Count ERROR Lines in a 10 GB Log File

Count ERROR lines in a giant log by streaming it with a generator expression in O(n) time and O(1) memory. Verified with a tracemalloc memory test.

The problem

lines is an iterator over a huge log file. Count the lines whose level (the first word) is exactly ERROR.

Same memory rule as the previous question: the test streams 300,000 lines and requires peak memory under 1 MB.

Examples

  1. Example 1

    Input

    count_errors(['INFO a\n', 'ERROR b\n', 'ERROR c\n', 'WARN d\n'])

    Expected output

    2
  2. Example 2

    Input

    check_streaming(count_errors, 300000)

    Expected output

    (60000, True)

+ 2 hidden tests on Submit — exact level only.

Edge cases to ask about

  • 'ERRORS' prefix
  • Lowercase 'error'
  • Empty file
How the tests call your code

These helpers run before your code. The test inputs above call them.

import tracemalloc as _tm

def log_stream(n):
    levels = ("INFO", "ERROR", "WARN", "INFO", "ERRORS")
    for i in range(n):
        yield f"{levels[i % 5]} message {i}\n"

def check_streaming(fn, n):
    _tm.start()
    try:
        count = fn(log_stream(n))
        peak = _tm.get_traced_memory()[1]
    finally:
        _tm.stop()
    return count, peak < 1024 * 1024

Hints

0/3

    How an interviewer scores this

    0/9
    Python 3.13 · count_errors
    ⌘/Ctrl + Enter runs the examples

    Your code runs in real CPython inside your browser — nothing is sent anywhere. The first run downloads the interpreter (about 6 MB, once). Your code is saved on this device as you type.

    Complexity Lab

    What does this cost as n grows?

    Interviewers score the analysis as much as the code. Commit to an answer first — then check it, and read why.

    Time complexity of the optimal solution
    Space complexity (extra memory)

    Pick both to reveal the answer.

    From brute force to optimal

    The progression an interviewer wants to hear, one step at a time.

    ApproachTimeSpaceIdea
    Generator expression + sumO(n)O(1)sum(1 for …) never builds a list.
    Walkthrough of the optimal approach (try it yourself first)

    sum(1 for line in lines if …) is a generator expression: it counts without building a list, so memory stays O(1). Match the level exactly — "ERRORS" and "error" must not count.

    At 10 GB the bottleneck is disk I/O. Mention reading in binary ('rb') and matching b"ERROR " to skip decoding, splitting the file into byte ranges for multiprocessing, or simply grep -c '^ERROR ', which is hard to beat.

    Complexity: O(n) time, O(1) space. One streaming pass with a single counter.

    Reveal the reference solution
    def count_errors(lines):
        return sum(1 for line in lines if line.startswith("ERROR ") or line.rstrip("\n") == "ERROR")

    Follow-ups interviewers ask

    • Count per hour from a timestamp.
    • Make it 4× faster on a 4-core machine.

    Frequently asked interview questions

    Core interview concepts, complexities, and follow-ups scored by hiring teams.

    What is the time complexity of Count ERROR Lines in a 10 GB Log File in Python?

    The optimal solution runs in O(n) time and O(1) auxiliary space. One streaming pass with a single counter.

    What follow-up questions do interviewers ask about Count ERROR Lines in a 10 GB Log File?

    Count per hour from a timestamp. Make it 4× faster on a 4-core machine.