When AI Writes Good Scala Code, and When It Doesn't

If your engineers use AI coding tools, you have probably seen both sides already. The same agent that writes a clean Scala service on Tuesday turns out to be a tangle of nulls and mutable state on Thursday, and the model did not change in between.

The difference comes from the codebase around the model. Scala AI tools now handle the language itself well, so whether their output is worth keeping depends on a short list of conditions that an experienced team controls. Get those conditions right, and AI speeds up your Scala work… Get them wrong, and it speeds up your technical debt.

Below, we walk through the four conditions where AI writes good Scala code, the four where it writes bad Scala code, and how the Scala compiler catches mistakes that slip through in a language like Python. We close with what all of it means for how you staff and review AI-assisted Scala work.

TL;DR

AI writes good Scala code when the domain types, the effect stack, and the compiler rules are already set and an experienced engineer reviews the output. AI writes bad Scala code when it starts from an empty repository, gets no stack decisions, or falls back to Java habits like null and var. Scala's compiler catches a whole class of AI mistakes at build time that plain Python typically surfaces at runtime, which puts senior Scala judgment at the center of any AI-assisted workflow.

Scala AI Tools Now Handle the Language Itself Well

Language support used to be the weak point for AI-generated Scala, and that is no longer where things break. Current models work comfortably across Scala 3, the major effect libraries, and the standard build tooling, and they can pick up an unfamiliar library by reading its documentation the way a new hire would.

That progress moves the failure point somewhere else. When AI-generated Scala goes wrong today, the cause is rarely syntax. The cause is missing context: no domain types to work against, no settled library choices, no compiler rules that reject bad output, and no one holding the architecture together. Our post on why Scala suits AI-assisted development covers the broader case for pairing the two. Here, the focus is the specific conditions that separate keepable AI output from rework.

When AI Writes Good Scala Code

AI writes good Scala code when the codebase already limits what correct code can look like. Four conditions do most of that work.

AI Writes Better Scala When the Domain Types Already Exist

Strong domain types turn an open-ended coding task into a fill-in-the-blanks task. When an experienced engineer has already defined the IDs, the states, and the error types, the AI writes an implementation against a signature that rules out whole categories of mistakes.

scala
opaque type CustomerId = String
opaque type InvoiceId  = String
 
enum InvoiceState:
  case Draft, Sent, Paid, Overdue
 
def markPaid(
  customer: CustomerId,
  invoice: InvoiceId
): IO[Either[BillingError, Invoice]]

With this signature in place, an AI agent cannot pass an invoice ID where a customer ID belongs or invent a status string like "PAID_LATE," because the compiler only accepts the four states the enum defines. The signature also spells out how a missing invoice gets reported, which leaves no natural place for a stray null. The types work as a specification the AI has to satisfy before the code compiles. These are some of the Scala language features that gain value once AI writes a share of your code.

AI Writes Better Scala When the Effect System Is Already Chosen

Effect systems like ZIO and Cats Effect give AI a narrow target. A function that returns IO[Either[BillingError, Invoice]] tells the model which business failures to expect and how to report them, so the model has far less room to improvise.

The word that matters is "chosen." One effect stack per codebase, declared up front and enforced in review, gives the AI a single consistent pattern instead of three competing ones.

AI Writes Better Scala When There Is an Existing Pattern to Copy

AI tools are strong pattern matchers. In a service-oriented codebase, the fifth service can follow the conventions of the first, and pointing an agent at an existing module to imitate produces far more consistent Scala than describing the conventions in prose. Teams with clean, business-focused types and a few well-built reference modules get noticeably better results from Scala AI tools than teams without them.

AI Writes Better Scala With Strict Compiler Settings

Instruction files like CLAUDE.md or AGENTS.md help, but models ignore written instructions some of the time, while a compiler rule fails the build every time it is broken. Teams that turn warnings into errors with -Werror, enable unused-code and discarded-value warnings, and run Scalafmt and Scalafix in CI give the AI an automatic first reviewer that rejects bad output before a human ever reads it. Reviewers should still watch for one workaround, an agent that quietly loosens the build settings or adds a @nowarn annotation to get past a failing check.

Experienced Scala teams treat build configuration as part of the AI setup, because every rule the compiler enforces is one less rule a reviewer has to catch by hand.

When AI Writes Bad Scala Code

AI writes bad Scala code when the repository leaves room for guessing. The failures cluster into four patterns, and each one gets more expensive the longer it stays in the codebase.

AI Writes Java-Style Scala Without Guardrails

Plenty of the Scala code available online is older "Scala as a better Java" code, written before modern functional libraries matured. Left unguided, AI models drift toward that style, and the result compiles, passes basic tests, and quietly works against everything Scala offers.

scala
// Java-flavored Scala an unguided model can produce
def lateFees(invoices: List[Invoice], today: LocalDate): BigDecimal = {
  var total = BigDecimal(0)
  for (i <- 0 until invoices.length) {
    if (invoices(i).state == "OVERDUE") {
      total = total + invoices(i).lateFee(today).get
    }
  }
  total
}
 
// Idiomatic Scala 3
def lateFees(invoices: List[Invoice], today: LocalDate): BigDecimal =
  invoices
    .filter(_.state == InvoiceState.Overdue)
    .flatMap(_.lateFee(today))
    .sum

The Java-flavored version keeps a mutable running total, indexes into a List one element at a time (which quietly makes the loop quadratic), compares the invoice state to a string, and calls .get on an optional fee, which throws the first time an overdue invoice has no fee attached. The idiomatic version does the job in four lines with no mutable state and nothing that can throw, and it makes an explicit choice to skip invoices with no fee instead of crashing on them. Scala's defaults around immutability exist to prevent exactly this kind of code, but only when the code actually uses them.

AI Writes Inconsistent Scala When It Starts From an Empty Repository

A blank repository gives AI nothing to imitate, so every session makes fresh decisions about naming, structure, and dependencies, and those decisions rarely agree with each other. The fix is to build the foundation by hand first, including the build, the stack, the domain model, and one complete feature from API to database, so the AI has a working pattern to follow from then on.

AI Mixes Scala Libraries When No Tech Stack Is Chosen

Scala has more than one credible answer for nearly every layer of a backend. Effects can run on ZIO, Cats Effect, plain Futures, or direct-style code, HTTP servers can run on http4s, Pekko HTTP, or ZIO HTTP with or without Tapir on top, and JSON has several solid libraries. An AI agent that receives no stack decision tends to mix them, sometimes within a single file. Choosing the stack for the AI, explicitly and early, removes a major source of messy output.

AI-Generated Scala Grows More Complex When Nobody Owns the Architecture

AI agents do their best work inside a single function or file and see much less of the system around it. When something breaks, the easiest move for an agent is to add code, whether that means another check, another wrapper, or another special case. Each addition makes sense on its own, and together they produce a codebase that grows faster than anyone can understand it.

Strong types slow this drift, but they cannot replace an engineer who owns the system design and pushes back when a change makes the code harder to reason about. That judgment is the part of the work AI has not taken over.

How the Scala Compiler Catches AI Mistakes That Python Lets Through

The strongest argument for Scala in AI-assisted development is what happens after the AI gets something wrong. In Python, many AI mistakes surface at runtime, sometimes weeks later in production. In Scala with strict compiler settings, many of the same mistakes fail the build.

scala
enum InvoiceState:
  case Draft, Sent, Paid, Overdue, Disputed
 
def nextStep(state: InvoiceState): String = state match
  case InvoiceState.Draft   => "send"
  case InvoiceState.Sent    => "wait for payment"
  case InvoiceState.Paid    => "close"
  case InvoiceState.Overdue => "send reminder"
 
// [warn] match may not be exhaustive.
// It would fail on pattern case: InvoiceState.Disputed

When a new Disputed state joins the invoice lifecycle and an AI-written function forgets to handle it, the Scala compiler flags the gap right away, and with -Werror enabled that warning blocks the build. The equivalent Python if-elif chain would fall through, return None, and keep running.

AI mistake What happens in Python What happens in Scala with strict compiler settings
Calls a library method that does not exist Fails with an AttributeError when that code path runs Fails to compile
Forgets to handle a new business state Falls through silently and returns None Fails the build with an exhaustiveness error
Swaps two IDs that share an underlying type Runs and writes bad data Fails to compile when the IDs are opaque types
Returns nothing for a missing record None travels through the code until something crashes An Option return type forces the caller to handle the missing case
Passes an argument of the wrong type Fails or misbehaves at runtime unless an optional type checker catches it Fails to compile

Python teams can close part of this gap with optional type checkers, but those checks sit outside the language and enforcement varies from codebase to codebase. Scala's checks run in the compiler on every build. For a team reviewing AI output, every mistake the Scala compiler catches is a mistake no human reviewer has to find.

Why AI Coding Tools Make Senior Scala Judgment More Valuable

AI changes the cost structure of a Scala team. Implementation time drops, while the time it takes to understand, review, and safely merge that implementation stays roughly where it was. The quality of AI-generated Scala ends up tracking the quality of the setup and review around it.

That shift changes where senior Scala experience pays off. The valuable work moves toward the tasks AI handles poorly:

  • Modeling the domain in types before any AI touches the implementation
  • Choosing one stack and enforcing it across the codebase
  • Configuring compiler flags, formatting, and linting so the build rejects bad output
  • Breaking AI-generated work into changes small enough to review properly
  • Owning the architecture and cutting complexity before it compounds

Junior engineers paired with AI and left without oversight carry more risk than before, because the volume of code they can produce now outruns their ability to evaluate it. Knowing what good Scala looks like in production still takes years of building and maintaining real systems.

If you work with an outside Scala team, ask how they constrain AI output. Ask which compiler flags run in CI, who owns the domain model, and how they keep AI-generated pull requests small enough to review. A team without clear answers ships faster today and hands you the cleanup later. Our guide to choosing a Scala outsourcing partner covers the rest of the vetting process.

AI writes good Scala when experienced engineers have already made the hard decisions, and bad Scala when it has to make those decisions on its own. Scala's type system and compiler give you a stronger safety net than dynamic languages offer, but someone still has to build that net and keep it tight.

Getting AI speed without AI cleanup takes Scala depth.

If your team is adding AI agents to a Scala codebase, a short conversation about types, compiler setup, and review flow can save months of rework. Talk to a Scala expert.

Frequently Asked Questions

Is AI good at writing Scala code?

Current AI models write solid Scala 3 code, including givens, type-level code, and effect-system code, when the codebase gives them clear types and patterns to follow. Output quality drops when the repository is new, the library stack is undecided, or no experienced engineer reviews the results.

Why does AI write Java-style Scala?

AI models learn from public code, and plenty of public Scala code follows an older "better Java" style with null values, mutable variables, and class-based data. Without guardrails, models drift toward that style. Domain types, strict compiler flags, and good reference code steer them toward idiomatic Scala.

How do you get AI to write idiomatic Scala?

Define the domain types and choose one effect stack before asking AI to write implementation code. Turn compiler warnings into errors with -Werror, run Scalafmt and Scalafix in CI, and point the AI at an existing module to imitate. Compiler-enforced rules work more reliably than written instructions in files like CLAUDE.md.

Is Scala a good language for AI-generated code?

Scala is a strong language for AI-generated code because its type system and compiler reject many AI mistakes at build time, including invented methods, unhandled states, and swapped IDs. The same mistakes in a dynamically typed language like Python often surface only at runtime. The benefit depends on a codebase that actually uses strong types and strict compiler settings.

How does using AI with Scala help outsourcing teams?

AI lets an experienced Scala team produce implementation code faster, which frees senior engineers to spend more time on domain modeling, architecture, and review. That gain only holds when the team constrains AI output with strong types, strict compiler settings, and review-sized changes. Buyers should ask any Scala vendor how they control the quality of AI-generated code.

Next
Next

Refinement Types in Scala 3: What You Can Use Today