PrepZone Logo
PrepZone

Batch Jobs for Large Data

Chunk-oriented processing, restartability and when to reach for Spring Batch.

Read these first

Why this matters

  • Loading a publisher's catalog CSV with 500,000 titles through a REST endpoint would time out and exhaust memory — batch jobs handle volume outside the request path.
  • Chunk-oriented processing commits after every N items: a crash at item 47,000 resumes at chunk 47, not item 1.
  • Spring Batch tracks job and step execution metadata in a database, giving you restartability, auditing, and skip/retry policies for poison records.
ControllerHTTP in/out — @RestController
ServiceBusiness rules — @Service
RepositoryDatabase — @Repository
DatabasePostgreSQL / H2
Each layer has one job. Dependencies point inward — controllers never touch the database directly.

Batch domain vocabulary

  • Job — the top-level unit of work, e.g. importPublisherCatalogJob.
  • Step — one phase of a job: read, process, write in chunks.
  • ItemReader — pulls input records one at a time (CSV, JDBC, JPA).
  • ItemProcessor — transforms or validates each item; return null to filter out.
  • ItemWriter — persists a chunk of processed items in one transaction.
  • JobRepository — stores execution state for restart and monitoring.

Dependencies and job metadata

Java
<dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-batch</artifactId>
</dependency>
Java
spring:
  batch:
    jdbc:
      initialize-schema: always  # creates BATCH_* tables; use Flyway in production
    job:
      enabled: false  # do not auto-run on startup; trigger via scheduler or API

CSV catalog import job

Java
@Configuration
public class CatalogImportBatchConfig {

    @Bean
    public Job importCatalogJob(JobRepository jobRepository,
                                Step importStep) {
        return new JobBuilder("importPublisherCatalog", jobRepository)
                .start(importStep)
                .build();
    }

    @Bean
    public Step importStep(JobRepository jobRepository,
                           PlatformTransactionManager txManager,
                           ItemReader<BookCsvRow> reader,
                           ItemProcessor<BookCsvRow, Book> processor,
                           ItemWriter<Book> writer) {
        return new StepBuilder("importStep", jobRepository)
                .<BookCsvRow, Book>chunk(100, txManager)
                .reader(reader)
                .processor(processor)
                .writer(writer)
                .faultTolerant()
                .skipLimit(50)
                .skip(ValidationException.class)
                .retryLimit(3)
                .retry(TransientDataAccessException.class)
                .build();
    }

    @Bean
    @StepScope
    public FlatFileItemReader<BookCsvRow> catalogReader(
            @Value("#{jobParameters['filePath']}") String filePath) {
        return new FlatFileItemReaderBuilder<BookCsvRow>()
                .name("catalogReader")
                .resource(new FileSystemResource(filePath))
                .delimited()
                .names("isbn", "title", "author", "price")
                .targetType(BookCsvRow.class)
                .build();
    }

    @Bean
    public ItemProcessor<BookCsvRow, Book> catalogProcessor() {
        return row -> {
            if (row.getPrice() <= 0) {
                throw new ValidationException("Invalid price for " + row.getIsbn());
            }
            return new Book(row.getIsbn(), row.getTitle(), row.getAuthor(), row.getPrice());
        };
    }

    @Bean
    public JpaItemWriter<Book> catalogWriter(EntityManagerFactory emf) {
        JpaItemWriter<Book> writer = new JpaItemWriter<>();
        writer.setEntityManagerFactory(emf);
        return writer;
    }
}

Trigger and monitor jobs

Java
@RestController
@RequestMapping("/api/admin/batch")
public class BatchJobController {

    private final JobLauncher jobLauncher;
    private final Job importCatalogJob;

    @PostMapping("/import-catalog")
    public ResponseEntity<String> triggerImport(@RequestParam String filePath) throws Exception {
        JobParameters params = new JobParametersBuilder()
                .addString("filePath", filePath)
                .addLong("timestamp", System.currentTimeMillis())
                .toJobParameters();

        JobExecution execution = jobLauncher.run(importCatalogJob, params);
        return ResponseEntity.accepted()
                .body("Job started: " + execution.getJobId());
    }
}
Java
@Scheduled(cron = "0 30 1 * * *")
public void nightlySalesAggregation(JobLauncher launcher, Job salesReportJob) throws Exception {
    launcher.run(salesReportJob, new JobParametersBuilder()
            .addLong("runDate", System.currentTimeMillis())
            .toJobParameters());
}

Quick recall

Everything you need if you only revisit this box.

  • Spring Batch is for high-volume, offline processing — not for request/response APIs.
  • Jobs are composed of steps; steps read-process-write in chunks with transactional commits per chunk.
  • JobRepository metadata enables restart from the last successful chunk after failure.
  • Configure skip and retry policies for rows that fail validation vs transient database errors.
  • Disable auto-start (spring.batch.job.enabled: false) and trigger jobs explicitly in production.

Test yourself

Answer these before moving on — recall is what makes it stick.