PrepZone Logo
PrepZone

The Stream API

Lazy pipelines, the collectors worth memorising, and an honest look at parallel streams.

Why this matters

  • Laziness is the single idea that explains why filter before map matters, why findFirst can stop early, and why a stream with no terminal operation does nothing at all.
  • The collectors are where most real stream work happens, and three or four of them cover nearly every case.
  • Parallel streams look like free speed and are usually not, and knowing exactly why is a frequent interview distinction.

The pipeline

Sourcelist.stream()
filterKeep some
mapTransform each
collectTerminal — runs everything
Nothing runs until the terminal operation. The intermediate steps only describe the work, which is why a pipeline without collect() or forEach() does nothing at all.
Java
List<String> result = orders.stream()                            // source
        .filter(order -> order.total().compareTo(MIN) > 0)        // intermediate — lazy
        .map(Order::customerName)                                 // intermediate — lazy
        .distinct()                                                // intermediate — lazy
        .sorted()                                                  // intermediate — lazy
        .limit(10)                                                 // intermediate — lazy
        .toList();                                                 // terminal — runs everything

Three properties that follow from laziness

  • Nothing runs without a terminal operation. A pipeline ending at map is dead code.
  • Elements flow one at a time, not stage by stage. Each element goes through filter, map and distinct before the next element starts.
  • Short-circuiting stops early. findFirst, anyMatch and limit stop consuming the source as soon as the answer is known.

The one-at-a-time behaviour is observable:

Java
List.of("a", "bb", "ccc").stream()
        .peek(value -> System.out.println("filter sees " + value))
        .filter(value -> value.length() > 1)
        .peek(value -> System.out.println("map sees " + value))
        .map(String::toUpperCase)
        .findFirst();
Java
filter sees a
filter sees bb
map sees bb

Only two elements were examined, and "ccc" was never touched. A stage-by-stage implementation would have processed all three twice.

Creating streams

Java
collection.stream();                             // from any Collection
Arrays.stream(array);                            // from an array
Stream.of("a", "b", "c");                        // from explicit values
Stream.empty();

IntStream.range(0, 10);                          // 0..9, primitive, no boxing
IntStream.rangeClosed(1, 10);                    // 1..10

Stream.iterate(1, n -> n * 2).limit(10);         // infinite, must be bounded
Stream.iterate(1, n -> n < 100, n -> n * 2);     // with a condition, Java 9+
Stream.generate(Math::random).limit(5);

Files.lines(path, UTF_8);                        // lazy over a file — close it
"a,b,c".chars();                                 // an IntStream of code points
Pattern.compile(",").splitAsStream("a,b,c");

Intermediate operations

Java
.filter(predicate)              // keep matching elements
.map(function)                   // transform each element
.flatMap(function)               // transform each into a stream, then flatten
.mapMulti(consumer)              // Java 16+, flatMap without allocating a stream per element
.distinct()                      // remove duplicates using equals
.sorted()                        // natural order
.sorted(comparator)
.limit(n)                        // at most n elements — short-circuits
.skip(n)                         // drop the first n
.takeWhile(predicate)            // Java 9+, stop at the first failure
.dropWhile(predicate)            // Java 9+, skip until the first failure
.peek(consumer)                  // observe without changing — for debugging only
.mapToInt(f) / .boxed()          // move between object and primitive streams

flatMap is the one worth a worked example, because the shape is not obvious:

Java
List<List<String>> nested = List.of(List.of("a", "b"), List.of("c"), List.of());

List<String> flat = nested.stream()
        .flatMap(List::stream)         // each inner list becomes a stream; all are concatenated
        .toList();                      // [a, b, c]

// A common real use: all distinct tags across all orders
Set<String> allTags = orders.stream()
        .flatMap(order -> order.tags().stream())
        .collect(Collectors.toSet());

Terminal operations

Java
.toList()                         // Java 16+, unmodifiable — the modern default
.collect(Collectors.toList())      // modifiable ArrayList
.collect(Collectors.toSet())
.forEach(consumer)                 // no ordering guarantee in parallel
.forEachOrdered(consumer)

.count()
.min(comparator) / .max(comparator)   // return Optional
.findFirst() / .findAny()              // return Optional; findAny is better in parallel
.anyMatch(p) / .allMatch(p) / .noneMatch(p)

.reduce(BinaryOperator)                      // Optional result
.reduce(identity, BinaryOperator)            // always a result
.sum() / .average() / .summaryStatistics()   // primitive streams only
.toArray(String[]::new)
Java
// allMatch and noneMatch on an empty stream are both true — vacuous truth
System.out.println(Stream.<String>empty().allMatch(s -> false));    // true

That surprises people and is mathematically correct: there is no element that fails the predicate.

Collectors

This is where most real work happens.

Java
// Counting by a key — the most used collector in practice
Map<String, Long> byDepartment = employees.stream()
        .collect(Collectors.groupingBy(Employee::department, Collectors.counting()));

// Grouping into lists
Map<String, List<Employee>> grouped = employees.stream()
        .collect(Collectors.groupingBy(Employee::department));

// Grouping with a downstream transformation
Map<String, List<String>> namesByDepartment = employees.stream()
        .collect(Collectors.groupingBy(Employee::department,
                 Collectors.mapping(Employee::name, Collectors.toList())));

// Summing and averaging per group
Map<String, Integer> payroll = employees.stream()
        .collect(Collectors.groupingBy(Employee::department,
                 Collectors.summingInt(Employee::salary)));

// Splitting into exactly two groups
Map<Boolean, List<Employee>> split = employees.stream()
        .collect(Collectors.partitioningBy(employee -> employee.salary() > 100_000));

// Building a map, with an explicit merge for duplicate keys
Map<String, Employee> byId = employees.stream()
        .collect(Collectors.toMap(Employee::id, employee -> employee,
                                  (existing, replacement) -> existing));

// Joining strings
String names = employees.stream()
        .map(Employee::name)
        .collect(Collectors.joining(", ", "[", "]"));

// Everything at once
IntSummaryStatistics stats = employees.stream()
        .collect(Collectors.summarizingInt(Employee::salary));
System.out.println(stats.getMax() + " " + stats.getAverage());

reduce

Java
// Sum, with an identity so the result is never Optional
int total = numbers.stream().reduce(0, Integer::sum);

// Without an identity, the result may be absent
Optional<Integer> product = numbers.stream().reduce((a, b) -> a * b);

// Prefer the primitive specialisation when it exists
int sum = numbers.stream().mapToInt(Integer::intValue).sum();

The identity must be a true identity for the operation — 0 for addition, 1 for multiplication, "" for concatenation. A wrong identity gives a wrong answer in parallel, because it is combined once per partition rather than once overall.

Primitive streams

Stream<Integer> boxes every element. The primitive specialisations do not.

Java
int sum = IntStream.rangeClosed(1, 1_000_000).sum();      // no boxing at all

OptionalDouble average = IntStream.of(1, 2, 3).average();

IntStream.of(1, 2, 3).boxed().toList();                   // back to objects when needed

// Mapping between them
employees.stream().mapToInt(Employee::salary).max();
IntStream.range(0, 5).mapToObj(i -> "item-" + i).toList();

For a million elements the difference is a million avoided allocations, which is a measurable win in a hot path.

Parallel streams, honestly

Java
long count = hugeList.parallelStream()
        .filter(this::expensiveCheck)
        .count();
AspectParallel helps whenParallel hurts when
Data sizeTens of thousands of elements or moreSmall collections — overhead dominates
Per-element costExpensive computationTrivial work like a field access
SourceArrayList or array — splits evenlyLinkedList — cannot split well
OperationsStateless and independentsorted, distinct, limit — need coordination
Work typeCPU-boundBlocking I/O — occupies the shared pool
OrderingOrder does not matterforEachOrdered serialises the output anyway
  • Data size

    Parallel helps whenTens of thousands of elements or more
    Parallel hurts whenSmall collections — overhead dominates
  • Per-element cost

    Parallel helps whenExpensive computation
    Parallel hurts whenTrivial work like a field access
  • Source

    Parallel helps whenArrayList or array — splits evenly
    Parallel hurts whenLinkedList — cannot split well
  • Operations

    Parallel helps whenStateless and independent
    Parallel hurts whensorted, distinct, limit — need coordination
  • Work type

    Parallel helps whenCPU-bound
    Parallel hurts whenBlocking I/O — occupies the shared pool
  • Ordering

    Parallel helps whenOrder does not matter
    Parallel hurts whenforEachOrdered serialises the output anyway

Parallel streams use the shared ForkJoinPool.commonPool(), so one slow pipeline affects the whole JVM.

Java
// Broken: shared mutable state, even with a synchronized list
List<String> results = Collections.synchronizedList(new ArrayList<>());
source.parallelStream().forEach(results::add);      // order lost, contention high

// Correct: let the collector handle combining
List<String> results = source.parallelStream().toList();

Common misreadings

  • "A stream is a collection." It holds no elements. It is a pipeline over a source.
  • "filter runs over everything, then map runs over the result." Each element traverses the whole pipeline before the next begins.
  • "A stream can be reused." It is single-use. A second terminal operation throws.
  • "peek is a good place for logging side effects." It may be skipped entirely. Use it only while debugging.
  • "parallelStream() makes things faster." Only for large, CPU-bound, independent work on a splittable source.
  • "Collectors.toList() is unmodifiable." It returns a modifiable ArrayList. .toList() is the unmodifiable one.
  • "toMap handles duplicate keys." It throws. Supply a merge function.
  • "findAny is non-deterministic in a sequential stream." In practice it returns the first element; the lack of a guarantee matters only in parallel.

Quick recall

Everything you need if you only revisit this box.

  • A stream is a lazy pipeline: intermediate operations build it, a terminal operation runs it, and nothing happens without one.
  • Elements flow one at a time through every stage, which is why short-circuiting works. Put filter and limit early.
  • Streams are single-use.
  • Know map vs flatMap (flatten nested structures), and takeWhile/dropWhile for ordered prefixes.
  • The collectors that matter: groupingBy (often with counting, mapping, summingInt), partitioningBy, toMap with a merge function, joining, summarizingInt.
  • reduce needs a true identity, or it breaks in parallel. Prefer IntStream.sum() where it exists.
  • Primitive streams avoid boxing — use mapToInt, boxed, mapToObj.
  • Parallel needs large data, expensive per-element work, a splittable source and no blocking. It shares commonPool() with the whole JVM.

Test yourself

Answer these before moving on — recall is what makes it stick.