Free Handbook · Every example compiled & verified

Interview Questions

60 Scala interview questions with model answers, from val and case classes to type classes, variance, Futures and Spark tuning, plus a coding round and take-home checklist.

0 / 142 lessons🔥 0 day streak
ShareXLinkedIn

Module 15 · what you'll be able to do

  • Answer the 25 junior Scala questions that decide first-round screens, precisely and with examples
  • Explain type classes, givens, variance, opaque types, Futures and effect systems at the depth a mid-level round expects
  • Reason through senior questions on Spark performance, streaming, codebase structure, migrations and production services
  • Run a live coding round with a repeatable script, and hand in a take-home that reviewers approve
01

How to use this module

Scala interviews follow a stable pattern: a screen on the language (immutability, case classes, pattern matching, Option, collections), a technical round on functional programming, types and concurrency, and, for experienced roles, a design conversation that assumes either data engineering with Spark and Kafka or backend services with an effect system. The 60 questions below are grouped the same way: 25 junior, 25 mid-level, 10 senior.

  • Say the answer out loud before opening it. Recognising an answer when you read it is not the same as producing it under pressure.
  • Read the "what they are really testing" line. It tells you what the interviewer will follow up on, which is where most candidates lose points.
  • Back every claim with code you have run. The modules these draw on, such as Case Classes & Pattern Matching, Generics & Givens and Futures & Concurrency, have examples you can paste into your editor.
Know which Scala they run
Many data teams are on Scala 2.12 or 2.13 because of Spark, while new backend services are usually Scala 3. Ask early, and when you answer, say which version a feature belongs to ("in Scala 3 that is a given; in 2.13 an implicit val"). It shows you can work in either codebase.
02

Junior: the language (10 questions)

Asked in nearly every first-round Scala screen. Short questions with precise answers; a vague answer here ends the interview early. Background: Module 01, Module 02 and Module 08.

JuniorWhat is the difference between val, var and def?

val is an immutable binding, evaluated once when defined. var is a mutable binding that can be reassigned. def defines a method, re-evaluated every time it is called. A lazy val is evaluated once, on first access. Note that val makes the reference immutable, not the object: val buf = ArrayBuffer(1) can still be changed with buf += 2. Idiomatic Scala uses val by default and var only for local, contained mutation.

What they are really testing: Evaluation timing, and that val does not make the object immutable.

JuniorWhy is "everything is an expression" important in Scala?

In Scala if, match, try and blocks all return values, so you write val label = if n > 0 then "positive" else "non-positive" instead of declaring a var and assigning it in branches. The value of a block is its last expression, and return is almost never needed. Statements that only have side effects return Unit, written (). This is what makes immutable code comfortable to write.

What they are really testing: That they write expression-oriented code rather than Java with different syntax.

JuniorWhat are Any, AnyVal, AnyRef, Nothing, Null and Unit?

Any is the top of the type hierarchy. AnyVal is the parent of value types (Int, Double, Boolean, Unit…), compiled to JVM primitives where possible. AnyRef is java.lang.Object, the parent of all classes. Nothing is the bottom type, a subtype of everything, with no values: throw has type Nothing, which is why val x: Int = if ok then 1 else throw Exception() type-checks, and why Nil is a List[Nothing] that fits any list. Null is the type of null. Unit has one value, ().

What they are really testing: Nothing as the bottom type, and why it makes throw and Nil fit anywhere.

JuniorWhat does == do in Scala compared with Java?

In Scala, == calls equals (after a null check), so it compares values: two strings with the same text are ==, and case classes compare field by field. Reference equality is eq. In Java, == on objects compares references, the source of many string bugs. Scala 3 also offers multiversal equality (import scala.language.strictEquality with derives CanEqual) to make comparing unrelated types, such as an Int with a String, a compile error.

What they are really testing: That == is value equality, and eq is reference equality.

JuniorWhat is string interpolation in Scala?

A string prefixed with an interpolator: s"Hello $name, you are ${age + 1}" inserts values; f"$price%.2f" adds printf-style formatting checked at compile time; raw"a\nb" keeps escapes literal. Interpolators are ordinary methods on StringContext, so libraries define their own: Doobie's sql"..." turns interpolated values into bound SQL parameters.

What they are really testing: The three standard interpolators and awareness that they are extensible.

JuniorHow does match differ from Java's switch?

match is an expression that returns a value, and it matches on patterns, not only constants: types (case s: String), case class structure (case Order(id, _, total)), lists (case head :: tail), tuples, guards (case n if n > 0) and alternatives (case "a" | "b"). There is no fall-through. A value that matches no case throws MatchError at run time, and for sealed types the compiler warns about missing cases.

What they are really testing: Patterns and guards, exhaustiveness warnings, and MatchError.

JuniorWhat is a for-comprehension?

Syntax for chaining map, flatMap and withFilter. for x <- xs; y <- ys if x != y yield (x, y) is rewritten by the compiler into xs.flatMap(x => ys.withFilter(y => x != y).map(y => (x, y))). So it works for any type with those methods, not only collections: Option, Either, Try, Future and effect types like IO. Without yield it becomes foreach.

What they are really testing: The desugaring, and that it works on Option, Either and Future.

JuniorWhat is Option and why use it instead of null?

Option[A] is either Some(value) or None. It puts "might be missing" into the type, so the compiler makes the caller deal with it, where null is invisible until it throws a NullPointerException. Work with it through map, flatMap, getOrElse, fold, pattern matching or a for-comprehension. Avoid .get, which throws on None. Wrap Java values that may be null with Option(javaValue).

What they are really testing: Avoiding .get, and wrapping nullable Java values.

JuniorWhat are higher-order functions? Give examples.

Functions that take functions as parameters or return them. The collection API is built on them: map, filter, flatMap, foldLeft, sortBy, groupBy. You write your own the same way: def retry[A](times: Int)(op: => A): A. Functions are values (val double: Int => Int = _ * 2) and compose with andThen and compose.

What they are really testing: Fluency with the collection HOFs and writing their own.

JuniorWhat is the difference between map and flatMap?

map applies a function to each element and keeps the structure: List(1, 2).map(n => List(n, n)) is List(List(1, 1), List(2, 2)). flatMap applies a function that returns a collection (or an Option, Future…) and flattens one level: the same call gives List(1, 1, 2, 2). On Option, flatMap chains steps that may each fail: parse(s).flatMap(validate). It is what makes for-comprehensions work.

What they are really testing: The flattening, and flatMap as "chain the next step that may fail".

03

Junior: classes, objects and traits (8 questions)

Interviewers check that you model data with case classes and sealed types, and know how Scala's object model differs from Java's. Background: Module 04 and Module 05.

JuniorWhat is a case class, and what does the compiler generate for it?

A class for immutable data. From case class Point(x: Int, y: Int) the compiler generates: public val fields, equals and hashCode based on the fields, a readable toString (Point(1,2)), a copy method with named arguments (p.copy(y = 5)), a companion apply (no new needed) and unapply for pattern matching. That makes case classes the default way to model data and messages.

What they are really testing: The full list, and copy for immutable updates.

JuniorWhat is a companion object?

An object with the same name as a class, in the same file. It holds what Java would make static: factory methods (apply), constants, and given instances for the class. A class and its companion can access each other's private members. The common pattern is a private constructor plus a companion apply or smart constructor that validates input and returns Either or Option.

What they are really testing: Replacing static, and the smart-constructor pattern.

JuniorWhat is a trait, and how does it differ from a Java interface or an abstract class?

A trait declares abstract and concrete members, including fields, and a class can mix in many of them: class Service extends Logging with Metrics. Unlike an abstract class, a class can extend several traits; since Scala 3 traits can also take parameters. When several traits override the same method, linearization decides the order, and super calls go to the next trait to the left, which enables stackable modifications. Use an abstract class when Java code must extend it or you need a constructor with Java interop.

What they are really testing: Multiple mixins, and a correct description of linearization.

JuniorWhat is a sealed trait, and why is it useful?

A sealed trait can only be extended in the same file, so the compiler knows every subtype. That makes pattern matching on it checkable: match warns "may not be exhaustive" if a case is missing, and adding a new subtype flags every match that must handle it. Sealed traits of case classes (or Scala 3 enums) model algebraic data types: PaymentStatus is exactly Pending, Paid(at) or Failed(reason).

What they are really testing: Exhaustiveness checking and modelling with ADTs.

JuniorWhat is an object in Scala?

A singleton: a class with exactly one instance, created lazily on first use and thread-safe. It is used for utility methods, constants, the program entry point (object Main { def main(args: Array[String]) }), companion objects and stateless implementations of a trait (case object for ADT cases without data). Scala 3 also allows top-level definitions, so many helpers no longer need a wrapping object.

What they are really testing: Singleton semantics and the main uses.

JuniorWhat are enums in Scala 3?

enum defines a closed set of cases, from simple ones (enum Color { case Red, Green, Blue }) to cases with parameters that form a full ADT (enum Shape { case Circle(r: Double); case Rect(w: Double, h: Double) }). Enums get values, valueOf and ordinal, support methods and derives, and are exhaustively checked in matches. They replace the Scala 2 idiom of a sealed trait with case objects.

What they are really testing: Scala 3 enums as ADTs, not only constants.

JuniorWhat is the difference between apply and unapply?

apply lets an object be called like a function: List(1, 2) is List.apply(1, 2), and a case class companion's apply builds an instance. unapply is the reverse, an extractor: it takes an object apart for pattern matching, returning an Option of its parts, so case Point(x, y) calls Point.unapply. You can write your own extractors, for example object Email { def unapply(s: String): Option[(String, String)] } to match strings as emails.

What they are really testing: Extractors as the mechanism behind pattern matching.

JuniorHow does Scala handle access modifiers?

Members are public by default (there is no public keyword). private restricts to the class and its companion; protected to the class and subclasses (stricter than Java: no package access). Qualified modifiers scope access: private[service] is visible in the service package, and private[this] (Scala 2) to the one instance. Constructor parameters without val are private to the class.

What they are really testing: Public by default and qualified access modifiers.

04

Junior: collections and errors (7 questions)

The collections API is where Scala developers spend their day, and handling failure as values is what separates idiomatic Scala from Java in Scala syntax. Background: Module 06 and Module 11.

JuniorWhat is the difference between immutable and mutable collections?

Immutable collections (List, Vector, Map, Set, the default imports) never change; operations return new collections that share structure with the old one, so they are cheap and safe to pass around and share between threads. Mutable collections (scala.collection.mutable.ArrayBuffer, mutable.Map) change in place. Use immutable by default, and mutable locally inside a function for performance, returning an immutable result.

What they are really testing: Structural sharing and the default-immutable convention.

JuniorWhen would you use List, Vector or Array?

List is a singly linked list: O(1) prepend and head/tail, O(n) indexing and append; ideal for recursion and pattern matching. Vector is a wide tree with effectively constant-time indexing, append and update; the best general-purpose immutable sequence. Array is a mutable JVM array: fastest for tight numeric loops and Java interop, but with Java's reference equality (Array(1) == Array(1) is false; use sameElements).

What they are really testing: Cost awareness, and the Array equality trap.

JuniorWhat does foldLeft do, and how is it different from reduce?

foldLeft(z)(op) starts with an initial value z and combines it with each element from left to right: List(1, 2, 3).foldLeft(0)(_ + _) is 6. The result type can differ from the element type, which makes it the general-purpose loop replacement (building a Map, counting, tracking state in a tuple). reduce has no initial value, uses the first element instead, so the result type must be the element type and it throws on an empty collection; reduceOption returns None instead.

What they are really testing: The initial value, and reduce failing on empty collections.

JuniorWhat are Either and Try?

Either[L, R] is Left(error) or Right(value); it is right-biased, so map and flatMap work on the Right and short-circuit on the first Left. Use it for expected failures with a typed error: Either[ValidationError, Order]. Try[A] is Success(value) or Failure(throwable), and Try { ... } catches non-fatal exceptions, which makes it the bridge from exception-throwing (often Java) code into values.

What they are really testing: When to choose each: typed domain errors versus captured exceptions.

JuniorHow do exceptions work in Scala?

Like Java's, but Scala has no checked exceptions: nothing forces you to declare or catch. try/catch/finally is an expression, and catch uses pattern matching (case e: IOException =>). Idiomatic Scala throws for bugs and truly exceptional situations, and returns Option, Either or Try for expected failures, so the possibility shows up in the type. scala.util.control.NonFatal catches everything except fatal errors like OutOfMemoryError.

What they are really testing: No checked exceptions, and errors as values for expected failures.

JuniorWhat is the difference between Nil, None, null, Nothing and Unit?

Nil is the empty List. None is the empty Option. null is the JVM null reference, to be avoided except at Java boundaries. Nothing is the bottom type with no values (the type of throw). Unit is the type of expressions with no useful value, with the single value (). The classic interview question, and all five appear in real code.

What they are really testing: Precision on five similar-sounding things.

JuniorWhat is a MatchError, and how do you avoid it?

A runtime exception thrown when a match expression receives a value that none of its cases handle: scala.MatchError: 7 (of class java.lang.Integer). Avoid it by matching on sealed types (the compiler then warns about missing cases, and -Werror turns the warning into a build failure), by adding a final case _ => when a default is truly correct, and by not using partial patterns in val definitions such as val List(a, b) = xs.

What they are really testing: Relying on the compiler's exhaustiveness check instead of catch-all cases.

05

Mid-level: functional programming (9 questions)

For roles with two to five years of experience. Scala teams expect you to be comfortable with functional ideas and to explain them in practical terms. Background: Module 07 and Module 09.

Mid-levelWhat is referential transparency, and why does it matter?

An expression is referentially transparent if you can replace it with its value without changing the program's behaviour. Pure functions (same input, same output, no side effects) give you that. It matters because such code can be reasoned about locally, tested without mocks, refactored safely, memoised and run in parallel. Side effects (I/O, mutation, time, randomness) break it; functional effect systems such as Cats Effect and ZIO describe effects as values (IO) so programs stay referentially transparent until the edge.

What they are really testing: The substitution definition and practical benefits, not only a slogan.

Mid-levelWhat are type classes, and how are they encoded in Scala 3?

A type class adds behaviour to types without inheritance: a trait with a type parameter (trait Show[A] { def show(a: A): String }), given instances for specific types (given Show[Int] with ...), and functions that require an instance through a using clause or context bound (def print[A: Show](a: A)). Extension methods make the syntax pleasant (a.show). The standard library uses them for Ordering and Numeric; Cats builds on Functor, Monad and Monoid; JSON libraries use Encoder/Decoder, often with derives.

What they are really testing: The three parts of the pattern and real library examples.

Mid-levelExplain given and using (implicits in Scala 2).

A given declares a canonical value of a type that the compiler can supply automatically; a using parameter asks the compiler to find one in scope. They replace Scala 2's implicit val and implicit parameters with clearer intent. Uses: type class instances, context passed through a call chain (an ExecutionContext, a transaction), and conversions (now explicit Conversion[A, B] instances). Givens are not imported by a wildcard; you write import x.given, which makes their origin visible.

What they are really testing: Knowing Scala 2 implicits and their Scala 3 replacements.

Mid-levelWhat is a monad, practically speaking?

A type F[A] with a way to put a value in (pure/Some/Right) and a flatMap that chains a computation depending on the previous result, obeying a few laws. Practically: Option chains steps that may be missing, Either steps that may fail with an error, Future steps that happen later, List steps with many results, and IO steps that perform effects. For-comprehensions are the syntax for monadic code. You do not need category theory to use them well.

What they are really testing: A practical definition with examples, not jargon.

Mid-levelWhat is the difference between strict and lazy evaluation in Scala?

Scala is strict by default: arguments and vals are evaluated immediately. Laziness is opt-in: lazy val (evaluated once on first access, thread-safe), by-name parameters (x: => A) (re-evaluated at each use, used by getOrElse, Option.fold and custom control structures), LazyList (elements computed on demand, allowing infinite sequences) and .view on collections (fusing map/filter without intermediate collections).

What they are really testing: The four opt-in lazy tools and when each is useful.

Mid-levelWhat is tail recursion, and what does @tailrec do?

A recursive call is in tail position when it is the last thing the function does; the compiler can then reuse the stack frame, turning the recursion into a loop that cannot overflow the stack. @tailrec does not enable the optimisation (it happens anyway for final methods); it makes the compiler fail if the function is not tail recursive, so a later edit cannot silently reintroduce a stack overflow. Non-tail recursion is usually made tail recursive with an accumulator parameter.

What they are really testing: That @tailrec is a check, and the accumulator technique.

Mid-levelWhat is currying, and where is it used?

A method with several parameter lists, def add(a: Int)(b: Int), can be applied one list at a time. Uses in practice: type inference flows from the first list to the next (foldLeft(0)(_ + _) knows the accumulator type from the first list); a final function parameter can be passed as a block (retry(3) { callApi() }), which makes custom control structures; and using clauses conventionally go in the last list. Partial application creates specialised functions.

What they are really testing: Inference and block syntax, the real reasons Scala code is curried.

Mid-levelWhat is the difference between a method and a function value?

A method (def) belongs to a class or object and can be generic, have multiple parameter lists, default arguments and using clauses. A function value is an object implementing Function1, Function2… (val f: Int => Int = _ + 1) that can be stored and passed around. The compiler converts a method to a function automatically when a function is expected (eta-expansion), which is why list.map(double) works with def double. Scala 3 also has polymorphic function types ([A] => List[A] => Int).

What they are really testing: Eta-expansion and what methods can do that functions cannot.

Mid-levelHow would you model a domain with ADTs to make illegal states unrepresentable?

Encode the rules in types instead of runtime checks. A payment is a sealed enum of Pending, Paid(at: Instant, ref: String) or Refunded(amount), not a class with five nullable fields and a status string. Use opaque types or validated wrappers for identifiers and quantities (opaque type Email = String with a smart constructor returning Either), Option only where absence is legal, and NonEmptyList where emptiness is not. The compiler then rejects code that builds an invalid state, and matches must handle every case.

What they are really testing: ADTs, opaque types and smart constructors used together.

06

Mid-level: the type system and Scala 3 (8 questions)

Scala's type system is its biggest strength and its biggest source of confusion. Interviewers want to see you use it to prevent bugs, not to show off. Background: Module 09.

Mid-levelExplain variance: covariance, contravariance and invariance.

Variance says how subtyping of a type parameter carries over to the generic type. Covariant List[+A]: a List[Cat] is a List[Animal], safe because immutable lists only produce values. Contravariant Function1[-A, +B]: a function accepting any Animal can be used where a function accepting Cat is needed. Invariant (default) Array[A] and mutable.Buffer[A]: both read and written, so neither direction is safe. The compiler enforces positions: a covariant type cannot appear as a method parameter, which is why List.prepend uses a lower bound [B >: A].

What they are really testing: Correct examples and why mutable collections must be invariant.

Mid-levelWhat are upper and lower type bounds?

An upper bound [A <: Animal] requires A to be a subtype of Animal, so its methods are available. A lower bound [B >: A] requires a supertype, used to let covariant collections accept wider elements: List[Cat].prepend(dog) returns a List[Animal]. Context bounds [A: Ordering] are different: they require a type class instance, not a subtype relationship.

What they are really testing: The lower-bound use case in covariant collections.

Mid-levelWhat are opaque types?

Scala 3 opaque type UserId = Long creates a distinct type that is a Long only inside its defining scope; outside, a UserId and a Long are incompatible, so you cannot pass an order ID where a user ID is expected. At run time it is a Long, with no wrapper object and no boxing cost. Expose a constructor (often validating) and extension methods in the companion. They replace most Scala 2 value classes.

What they are really testing: Type safety with zero runtime cost.

Mid-levelWhat are extension methods, and how do they relate to Scala 2 implicit classes?

Scala 3 extension (s: String) def isEmail: Boolean = ... adds a method to an existing type without changing it or wrapping it. They are how type class syntax is provided (a.show) and how libraries enrich Java types. Scala 2 used implicit class RichString(s: String), which relied on an implicit conversion; extension methods are more direct and easier to discover. They are resolved only when in scope, so you import them deliberately.

What they are really testing: The Scala 2 to 3 migration story.

Mid-levelWhat does derives do?

case class User(name: String, age: Int) derives Codec, CanEqual asks the compiler to generate given instances of those type classes for the type, using each type class's derived method, which inspects the type's structure at compile time (Scala 3 Mirrors and inline macros). Libraries like Circe, uPickle, ZIO JSON and Tapir support it, removing the boilerplate of writing encoders and decoders by hand.

What they are really testing: Automatic, compile-time type class derivation.

Mid-levelWhat is type erasure, and how does it affect Scala code?

On the JVM, generic type arguments are erased at run time: a List[Int] and a List[String] are both just List. So case xs: List[Int] => cannot check the element type; the compiler warns that the pattern is unchecked. Work around it by matching on the elements, using sealed ADTs instead of type tests, or requiring a ClassTag or TypeTest when a runtime type check is genuinely needed (for example creating an Array[A]).

What they are really testing: The unchecked-pattern warning and ClassTag.

Mid-levelWhat are the main differences between Scala 2 and Scala 3?

Syntax: optional braces with indentation, if ... then, @main methods and top-level definitions. Contextual abstractions: given/using, extension methods and explicit Conversion instead of the single implicit keyword. Types: enums, opaque types, union (A | B) and intersection (A & B) types, match types, trait parameters, and derives. Metaprogramming: new inline and quote-based macros (Scala 2 macros do not carry over). Scala 3 can use most Scala 2.13 libraries, and many Spark codebases still run on 2.12 or 2.13.

What they are really testing: Concrete differences and awareness that Spark keeps teams on Scala 2.

Mid-levelWhat are union and intersection types?

A union type A | B is a value that is either an A or a B, without a common wrapper: def parse(s: String): Int | ParseError. Match on it with type patterns. An intersection type A & B is a value that is both, for example a parameter that must be Closeable & Flushable. Unions are useful for ad-hoc error types and Java interop; for domain modelling a sealed ADT is usually clearer.

What they are really testing: When a union is appropriate versus a sealed ADT.

07

Mid-level: concurrency, the JVM and tooling (8 questions)

How Scala runs: Futures and thread pools, effect systems, Java interop, builds and tests. Background: Module 10 and Module 12.

Mid-levelHow does Future work, and what is an ExecutionContext?

A Future[A] represents a value that will be available later; Future { work } schedules the work on a thread pool and returns immediately. The ExecutionContext is that thread pool, passed as a using parameter to Future.apply, map, flatMap and callbacks. Futures are eager (they start as soon as they are created) and memoised (the result is computed once). Compose them with map/flatMap/for, handle failures with recover, and never block with Await except at the program's edge.

What they are really testing: Eagerness and the role of the ExecutionContext.

Mid-levelWhy do these two for-comprehensions over Futures behave differently?

In for a <- Future(slowA()); b <- Future(slowB()) yield a + b, the second Future is created inside flatMap, only after the first completes, so the calls run sequentially. Creating them first, val fa = Future(slowA()); val fb = Future(slowB()); for a <- fa; b <- fb yield a + b, starts both immediately, so they run in parallel. The difference comes from Futures being eager. With lazy effect types (IO), you use parMapN or par combinators instead.

What they are really testing: Understanding eagerness through a classic bug.

Mid-levelWhy should blocking calls not run on the global ExecutionContext?

The global pool has about one thread per CPU core. A blocking call (JDBC, a file read, Thread.sleep, a synchronous HTTP client) holds a thread doing nothing, and a few of them can starve the pool so every other Future in the application stalls. Run blocking work on a separate, larger pool dedicated to blocking (ExecutionContext.fromExecutor(Executors.newFixedThreadPool(n))), or wrap it in blocking { } so the global pool can add a thread. Effect systems provide IO.blocking. On Java 21+, virtual threads are another option.

What they are really testing: Thread-pool starvation and how to isolate blocking work.

Mid-levelWhat are Cats Effect and ZIO, and why would a team choose them over Future?

Both are effect systems: an IO (or ZIO) value describes a computation without running it, so it is lazy and referentially transparent, retryable and composable. They add what Future lacks: cancellation and interruption, resource safety (Resource, Scope), structured concurrency with lightweight fibers, timeouts, retries and scheduling, plus ecosystems for HTTP (http4s, zio-http), streaming (fs2, ZStream) and databases. Teams choose them for large concurrent services; the cost is a learning curve and a different programming style.

What they are really testing: Concrete capabilities, not only "it is more functional".

Mid-levelHow do you share mutable state safely between threads in Scala?

First, avoid it: prefer immutable data and passing messages or results. When shared state is needed: java.util.concurrent.atomic types (AtomicLong, AtomicReference with updateAndGet) for single values; concurrent collections (ConcurrentHashMap, TrieMap); synchronized blocks for small critical sections; and in effect systems, Ref (atomic, functional) and Semaphore. @volatile only guarantees visibility, not atomic read-modify-write. The classic bug is var count = 0 incremented from several Futures.

What they are really testing: Atomic updates versus visibility, and the lost-update bug.

Mid-levelHow does Scala interoperate with Java?

Scala compiles to JVM bytecode and can call any Java library directly. Details to handle: Java values may be null (wrap with Option(x)); collections need conversion (import scala.jdk.CollectionConverters.*, then asScala/asJava); Java functional interfaces accept Scala lambdas automatically (SAM conversion); checked exceptions are not enforced; java.util.Optional converts with scala.jdk.OptionConverters. Calling Scala from Java is harder (default arguments, givens, companion objects), so APIs meant for Java users are designed with that in mind.

What they are really testing: Null, collection conversion and SAM conversion.

Mid-levelWhat is sbt, and how do you structure a multi-module build?

sbt is Scala's build tool: build.sbt declares projects, settings and dependencies (%% adds the Scala version suffix to Scala libraries), and the sbt shell runs tasks such as compile, test and ~testQuick incrementally. A multi-module build defines several projects with lazy val core = project, lazy val api = project.dependsOn(core), and an aggregating root. Common settings go in ThisBuild or a shared settings sequence. Plugins (sbt-assembly, sbt-native-packager, scalafmt, sbt-tpolecat) go in project/plugins.sbt.

What they are really testing: Practical build knowledge, including %% and dependsOn.

Mid-levelHow do you test Scala code?

Unit tests with MUnit or ScalaTest (both run by sbt test); property-based tests with ScalaCheck (generate hundreds of random inputs and check an invariant, such as "sorting twice equals sorting once"); for effectful code, munit-cats-effect or zio-test. Keep business logic pure so most tests need no mocks, and put I/O behind traits with fake implementations. Integration tests against real dependencies use Testcontainers (Postgres, Kafka). For Spark, test transformations on small local DataFrames with a local SparkSession.

What they are really testing: Property-based testing and testing pure code without mocks.

08

Senior: data pipelines, systems and production (10 questions)

Senior Scala interviews assume the ecosystem: Spark and Kafka for data roles, an effect system and a JVM service in production for backend roles. There is no single right answer; the model answers show the shape of a strong one: what you would ask first, what you would measure, and which trade-off you would accept.

SeniorA Spark job written in Scala is slow. How do you diagnose and fix it?

Start with the Spark UI: find the slow stage, then look at task-time distribution (skew shows as a few tasks far slower than the median), shuffle read/write sizes and spill to disk. Common fixes: avoid unnecessary shuffles (filter and select columns early, reduceByKey/aggregations instead of groupByKey); fix skew by salting hot keys or relying on AQE skew-join handling; broadcast small tables in joins; right-size partitions (spark.sql.shuffle.partitions, AQE coalescing); prefer DataFrame/Dataset operations and built-in functions over UDFs and typed lambdas that block Catalyst optimisation; cache only what is reused; and read columnar formats (Parquet, Delta, Iceberg) with partition pruning. Measure after each change.

What they are really testing: A systematic, UI-driven approach and knowledge of skew, shuffles and broadcast joins.

SeniorHow would you design a streaming pipeline with exactly-once semantics?

True end-to-end exactly-once needs cooperation from every part. With Kafka as the source: consume in a consumer group, and either use Kafka transactions (read-process-write to Kafka with isolation.level=read_committed, as Kafka Streams and Flink do), or make the sink idempotent (upserts keyed by a natural ID, or storing the processed offset in the same database transaction as the result). Spark Structured Streaming achieves it with checkpointed offsets plus idempotent or transactional sinks (Delta Lake). Plan for replay, late and out-of-order events (watermarks), schema evolution (a schema registry) and back-pressure.

What they are really testing: That exactly-once is a property of the whole pipeline, and idempotent sinks.

SeniorHow do you structure a large Scala codebase so that it stays maintainable?

Split it into sbt modules along domain boundaries, with a small core of pure domain code (ADTs, business rules) that has no dependency on HTTP, databases or frameworks, and adapters around it (the "functional core, imperative shell" or ports-and-adapters approach). Agree on one effect style (Future, Cats Effect or ZIO) and one JSON and HTTP stack instead of mixing. Enforce conventions with scalafmt, scalafix, strict compiler flags (-Werror, unused warnings, sbt-tpolecat) and code review. Keep implicits and type-level tricks to where they pay off, because the team's ability to read the code is the real constraint.

What they are really testing: Pragmatic restraint with the language's power, and enforced conventions.

SeniorFuture, Cats Effect or ZIO: how do you decide for a new service?

Consider the team and the problem. Future is in the standard library and fine for simple services and Play or Pekko-based systems, but lacks cancellation and resource safety. Cats Effect (with http4s, fs2, Doobie) suits teams comfortable with type-class-based functional programming and wants the Typelevel ecosystem. ZIO offers a batteries-included, more concrete API (typed errors, environment, layers) that many find easier to onboard. Hiring, existing code, library needs and operational tooling matter more than benchmarks. Decide once and document why.

What they are really testing: Decision criteria grounded in team and ecosystem.

SeniorHow do you handle schema evolution for data produced and consumed by Scala services?

Use a schema format with explicit compatibility rules: Avro or Protobuf with a schema registry, or versioned JSON with a documented contract. Evolve with backward- and forward-compatible changes only: add optional fields with defaults, never reuse or retype a field, deprecate before removing. Consumers must tolerate unknown fields. For data lakes, table formats like Delta and Iceberg support additive schema changes and time travel. Generate Scala case classes from schemas (or derive codecs), and add compatibility checks to CI so a breaking change fails the build before it reaches production.

What they are really testing: Compatibility rules and automated enforcement.

SeniorHow would you migrate a codebase from Scala 2.13 to Scala 3?

Check the blockers first: Scala 2 macros (and libraries that depend on them) do not work in Scala 3, and Spark jobs may need to stay on 2.13. Move to the latest 2.13 with -Xsource:3 and fix the warnings, update dependencies to versions published for Scala 3, and use scalafix rules and the Scala 3 migration tooling. Scala 3 can consume 2.13 libraries (and vice versa via TASTy), so modules can migrate one at a time with cross-building. Replace implicits with given/using and extension methods gradually after the compile succeeds, not during the migration.

What they are really testing: The macro blocker, cross-building and incremental migration.

SeniorHow do you make a JVM-based Scala service production-ready?

Observability: structured JSON logging (logback or log4cats) with correlation IDs, metrics exported to Prometheus, and distributed tracing with OpenTelemetry. Health and readiness endpoints. Configuration from the environment with typed loading (PureConfig, Ciris) and secrets from a vault. JVM settings: heap sized from the container limit (-XX:MaxRAMPercentage), a suitable GC (G1 by default, ZGC for low latency), and GC logs. Graceful shutdown that stops accepting requests and drains in-flight work. Timeouts, retries with back-off and circuit breakers on every outbound call. Load-test with realistic data before launch.

What they are really testing: Operational maturity, including JVM specifics in containers.

SeniorHow do you keep compile times manageable in a large Scala project?

Measure first (sbt's -Vstatistics-style profiling, or the scalac profiler) to find hot spots. Common causes: heavy implicit search and derivation (for example deriving codecs for large nested ADTs in many places; derive once in companions instead), large files that recompile together, broad wildcard imports, and deep module dependency chains. Split into smaller sbt modules so incremental compilation touches less, keep the dependency graph shallow, avoid unnecessary macros, use a build server (Bloop, or sbt's own server) and remote caching in CI. Fast feedback is a team productivity issue.

What they are really testing: Knowing the typical causes, especially implicit derivation.

SeniorDesign a rate limiter for an API written in Scala.

Clarify: per user, per IP or per API key; a single instance or a cluster; hard limit or burst allowance. For one instance, a token bucket per key held in a ConcurrentHashMap or a Cats Effect Ref, refilled based on elapsed time, allows bursts up to the bucket size. For a cluster, keep state in Redis with an atomic Lua script or the sliding-window-log/counter approach, and accept small inaccuracies for speed. Return 429 with a Retry-After header, expose metrics on rejections, expire idle keys to bound memory, and decide what happens when Redis is unavailable (fail open or closed).

What they are really testing: Clarifying questions, algorithm choice and failure modes.

SeniorWalk through how you review a Scala pull request.

Correctness first: edge cases, error handling (are failures values, are they handled or silently dropped?), exhaustive matches, no .get or head on things that can be empty. Types: do they express the domain (ADTs, opaque types) or leak primitives and nulls? Effects and concurrency: blocking calls on the right pool, no shared mutable state, resources released. Readability: implicits and advanced type features used only where they pay off, names that explain intent, no clever point-free code the team cannot maintain. Tests cover the behaviour, including a property test where an invariant exists. Performance: collection choice, no accidental O(n²), and for Spark, no unnecessary shuffles or collects. Comments are specific and separate blocking issues from suggestions.

What they are really testing: A structured review, with Scala-specific failure modes.

The senior-answer shape
Clarify the goal and constraints, name two options with their costs, pick one and say what would make you change your mind, then say how you would verify it in production. That structure matters more than any single fact.
09

The coding round, walked through

A live coding round is 30 to 45 minutes on one or two problems in a shared editor. The interviewer grades how you think, communicate and test, not only whether it compiles. Follow the same script every time, the one from Module 14.

  1. 1
    Clarify (2 min)

    Restate the problem. Ask about input size, empty input, duplicates, ordering of the output, and what to return when there is no answer (an Option?). Write the signature first.

  2. 2
    Example (1 min)

    Work one small case by hand. It becomes your first test.

  3. 3
    Brute force out loud (2 min)

    "For every request, count the same client's requests in the next window: O(n²)." Say it and its cost before improving it.

  4. 4
    Pick the pattern (1 min)

    groupBy then a window? A heap? Name it, and why.

  5. 5
    Code (15 min)

    Talk while you type. Immutable collections and a fold by default; a local mutable loop is fine if you say why.

  6. 6
    Test (5 min)

    Run your example, then the edges. Finding your own bug scores higher than never having one.

  7. 7
    Complexity and improvements (2 min)

    State time and space, then the better version if there is one.

A typical 30-minute problem solved that way: given a log of (client, second) requests in no particular order, return every client that made more than limit requests within any window-second span, sorted by name. The version below groups by client, sorts each client's times, and slides a window over them (Module 14, pattern 2).

scalaMain.scala
// Contract: log is unsorted and may be empty; limit >= 1; window >= 1 second.
// A client offends if more than `limit` requests fall within `window` seconds
// (times t and t + window - 1 are in the same span). Result sorted by name.
def offenders(log: Seq[(String, Int)], limit: Int, window: Int): List[String] =
  log
    .groupMap(_._1)(_._2)
    .collect {
      case (client, times) if exceeds(times.sorted, limit, window) => client
    }
    .toList
    .sorted

def exceeds(sorted: Seq[Int], limit: Int, window: Int): Boolean =
  val v = sorted.toVector
  v.indices.foldLeft((0, false)) { case ((left, found), right) =>
    if found then (left, true)
    else
      val newLeft = (left to right).find(l => v(right) - v(l) < window).getOrElse(right)
      (newLeft, right - newLeft + 1 > limit)
  }._2

@main def run(): Unit =
  val log = Seq(
    "ana" -> 1, "bo" -> 2, "ana" -> 3, "ana" -> 5, "bo" -> 30, "cy" -> 7,
    "ana" -> 70, "bo" -> 31, "bo" -> 33, "cy" -> 50, "bo" -> 32,
  )
  println(offenders(log, limit = 2, window = 10).mkString(", "))
  println(offenders(log, limit = 3, window = 10).mkString(", "))
  println(offenders(Seq.empty, limit = 1, window = 1).size) // empty log
Outputcompiled & run with real Scala
ana, bo
bo
0

ana makes 3 requests in seconds 1 to 5, so she offends at a limit of 2 but not 3; bo makes 4 in seconds 30 to 33. The left edge only moves forward, so each client's window scan is O(k); overall O(n log n) for the sorts.

Your turn

The interviewer follows up: "the log is now an endless stream, already in time order". Rewrite it with a Map[String, Queue[Int]] of each client's recent times, dropping times that fall out of the window as each request arrives. What is the memory cost now?

What loses the round

  • Silence for ten minutes, then a wall of code
  • Printing a groupBy result and assuming its order
  • Calling .head or .get on something that can be empty
  • An elaborate type-level solution the interviewer cannot follow
  • "It should work" without running an example

What wins it

  • Narrating your reasoning, including dead ends
  • A signature, contract and example before code
  • Testing edge cases: empty log, a limit no one exceeds, requests exactly window seconds apart
  • Naming the complexity without being asked
  • Idiomatic, readable Scala: a fold, a pattern match, an Option
10

Take-home assignment checklist

Scala take-homes are usually "build a small service" (often with http4s, ZIO or Play) or "process this dataset" (plain Scala or Spark). Reviewers open the README, run sbt test, read the tests, then the code. Most rejected submissions fail at step two: it does not build on the reviewer's machine.

  • It builds on a clean machine: sbt test works with only a JDK and sbt installed. Pin the sbt version in project/build.properties and the Scala version in build.sbt.
  • README: what it does, how to build, run and test it in three commands, example requests or sample input and output, and the decisions you made, including what you deliberately left out.
  • Tests with MUnit or ScalaTest: the happy path, empty and invalid input, and the one tricky rule in the spec. A ScalaCheck property where an invariant exists stands out.
  • Clean compile: no warnings, with strict flags (sbt-tpolecat or -Werror), and code formatted with scalafmt.
  • Idiomatic types: case classes and sealed ADTs for the domain, Option and Either for absence and errors, no null, no .get, no var leaking out of a function.
  • Structure: pure domain logic separated from I/O, so the core is tested without a server or a database.
  • Errors on purpose: invalid input returns a clear error (400 with a message for an API, a readable message for a CLI), not a stack trace.
  • Consistency: one effect style (Future, Cats Effect or ZIO), not a mix, and libraries the company uses if the brief mentions them.
  • Time-box to what they asked (usually 3 to 4 hours) and say so in the README. A handful of meaningful commits beats one "final" dump.
The sentence reviewers want to write
"Built first time, the types model the domain, the tests cover the edge cases, and the code reads like the team already wrote it." Aim every decision at that sentence. The next module, Job Ready, turns the same standards into a portfolio.

Frequently asked questions

What Scala topics are asked most in interviews?
At junior level: val versus var, expressions, case classes, pattern matching, sealed traits, Option, Either and Try, collections and folds, and for-comprehensions. At mid level: type classes and givens, variance and bounds, opaque types, Scala 2 versus 3, Futures and execution contexts, effect systems and testing. Senior rounds add Spark performance, streaming, codebase structure and production concerns.
Do Scala interviews expect Spark knowledge?
For data-engineering roles, yes: expect questions on DataFrames and Datasets, transformations versus actions, shuffles, partitioning, skew and joins, often alongside Kafka. Backend Scala roles focus instead on functional programming, an effect system (Cats Effect or ZIO) and HTTP services.
Should I prepare Scala 2 or Scala 3?
Learn Scala 3, but know the Scala 2 equivalents: implicits for givens and using, implicit classes for extension methods, sealed traits with case objects for enums. Many Spark codebases are still on Scala 2.12 or 2.13, so being able to read both is valuable.

Finish the Scala handbook, then get hired

Sit the exam for your certificate, run your resume through the ATS checker, and see the jobs that ask for exactly this.

Check my resume
Found this course useful? Share it.
ShareXLinkedIn

Comments

0

Join the conversation. Sign in to leave a comment — we'd love to hear your thoughts.