L2 · Working engineerAdvanced L2~8 min · 3 tests#93

Parallel CPU Work With multiprocessing

Square numbers in parallel with multiprocessing.Pool or ProcessPoolExecutor and learn why CPU-bound work needs processes, not threads, because of the GIL.

The problem

Write parallel_squares(nums) that squares every number using a pool of processes — multiprocessing.Pool or concurrent.futures.ProcessPoolExecutor — and returns the results in input order.

Define the worker function at module level (pools pickle it to send it to other processes).

Running in your browser: Python here runs on WebAssembly, which has no OS threads or processes. The standard APIs still work — threading, concurrent.futures, multiprocessing, asyncio — but they run on a deterministic simulator: threads run to completion when started, pools run tasks in order, and asyncio uses a virtual clock (await asyncio.sleep(0.2) advances time by 0.2 s instantly). Write exactly the code you would write in the interview.

Examples

  1. Example 1

    Input

    used_pool(parallel_squares, [1, 2, 3, 4, 5])

    Expected output

    ([1, 4, 9, 16, 25], True)

+ 2 hidden tests on Submit.

Edge cases to ask about

  • Empty input
  • Order of results
  • Lambdas cannot be pickled
How the tests call your code

These helpers run before your code. The test inputs above call them.

def used_pool(fn, nums):
    IV_POOL_USES[0] = 0
    out = fn(nums)
    return list(out), IV_POOL_USES[0] > 0

Hints

0/3

    How an interviewer scores this

    0/9
    Python 3.13 · parallel_squares
    ⌘/Ctrl + Enter runs the examples

    Your code runs in real CPython inside your browser — nothing is sent anywhere. The first run downloads the interpreter (about 6 MB, once). Your code is saved on this device as you type.

    Complexity Lab

    What does this cost as n grows?

    Interviewers score the analysis as much as the code. Commit to an answer first — then check it, and read why.

    Time complexity of the optimal solution
    Space complexity (extra memory)

    Pick both to reveal the answer.

    From brute force to optimal

    The progression an interviewer wants to hear, one step at a time.

    ApproachTimeSpaceIdea
    ThreadsO(n)O(n)The GIL lets only one thread run Python bytecode at a time — no speed-up for CPU work.
    bestProcess poolO(n / cores)O(n)Separate interpreters, separate GILs; results are pickled back.
    Walkthrough of the optimal approach (try it yourself first)

    CPU-bound pure-Python code does not speed up with threads: the GIL lets only one thread execute bytecode at a time. Processes each have their own interpreter and GIL, so ProcessPoolExecutor().map(square, nums) uses every core and returns results in order.

    Costs to mention: process start-up, and pickling arguments and results. For tiny tasks like squaring, the overhead dominates — use chunksize, or NumPy, which releases the GIL. On Windows/macOS, guard the entry point with if __name__ == "__main__":.

    Complexity: O(n) time, O(n) space. The total work is n squarings; with p processes the wall-clock time is about n / p plus the cost of pickling data between processes.

    Reveal the reference solution
    from concurrent.futures import ProcessPoolExecutor
    
    
    def square(x):
        return x * x
    
    
    def parallel_squares(nums):
        with ProcessPoolExecutor() as pool:
            return list(pool.map(square, nums))

    Follow-ups interviewers ask

    • Why is this slower than a list comprehension for 5 numbers?
    • Threads vs processes vs asyncio — when to use which?

    Frequently asked interview questions

    Core interview concepts, complexities, and follow-ups scored by hiring teams.

    What is the time complexity of Parallel CPU Work With multiprocessing in Python?

    The optimal solution runs in O(n) time and O(n) auxiliary space. The total work is n squarings; with p processes the wall-clock time is about n / p plus the cost of pickling data between processes.

    What is the brute-force approach, and how do you optimise it?

    Threads: O(n) time, O(n) space. The GIL lets only one thread run Python bytecode at a time — no speed-up for CPU work. Process pool: O(n / cores) time, O(n) space. Separate interpreters, separate GILs; results are pickled back.

    What follow-up questions do interviewers ask about Parallel CPU Work With multiprocessing?

    Why is this slower than a list comprehension for 5 numbers? Threads vs processes vs asyncio — when to use which?