Matchbox: from 12 GB to 4 GB – 24 hours with an AI agent and a test setup

26.09.26 | Oliver Egger

In 24 hours, Matchbox went from a fixed 12 GB heap to a 4 GB container for a typical validation server. The work was done by an AI agent, Claude Code with Opus 5.5: it ran the builds, the load tests and the heap dump analysis, and proposed and implemented the optimizations, while a human stayed in the loop for every decision. This article describes how we got there, and why a reproducible load test mattered as much as the agent.

Matchbox and why validation matters

FHIR®, the HL7 standard for healthcare interoperability, leaves a lot of freedom in how data is represented. To guarantee the correctness of an exchange like a lab report or a patient summary, FHIR resources are validated against the rules published in FHIR implementation guides (IGs): which fields are required, which codes are allowed, which cardinalities apply.

Matchbox is our open-source FHIR validation and mapping tooling. It wraps the official HL7 FHIR validator in a server with its own FHIR API and a web GUI, supports multiple IGs at once, and also offers validation through an MCP service. We started in August 2020 and have published 98 releases since then. It runs in production infrastructure (e.g. the Swiss notification of laboratory results, ch-elm), in test platforms such as IHE Gazelle at IHE connectathons and projectathons, and in integration pipelines.

Memory became the bottleneck

FHIR IGs tend to get bigger: to stay aligned with international developments, they depend on European and international IGs, and each new terminology dependency makes validation hungrier for memory. Initially, we assumed this was a natural evolution and let the container grow. Then we were asked why an existing validation setup needed four times the memory within one year. In addition, several issues and pull requests from the community pointed to memory problems:

  • Achraf (@achrafachkari) found that the human-readable HTML narrative of every loaded resource stayed in memory (#566).

  • Valentin (@reva) measured better throughput with a smaller heap than the image's default of 12 GB and asked for configurable JVM options (#594, #595).

  • We ourselves hit an OutOfMemoryError on our memory-constrained test instance, while the health check still reported the server as up (#457).

In release 4.1.14 we had already added memory metrics to Matchbox, but observing the problem was not enough. Fortunately, we also had test setups for load testing.

The test setup: reproduce before you optimize

The test setups were mostly used to reproduce concurrency issues in validation: Apache JMeter sends 8,000 validations of a ch-elm lab report (4 threads) to a Matchbox container and asks the server for its memory use after each one. The results were trustworthy, but the comparison was manual on my development machine, and reruns were difficult to compare because the machine load could differ between runs.

So we first made the measurements reproducible. A script starts a fresh container and records the startup time, the first and the second validation, and the live heap. All runs happened on the same idle machine. To check servers with several IGs, a second JMeter test validates examples from different IGs, each of which gets its own validation engine. For a validator, a faster answer is worthless if it is a different answer, so every run also compared the validation results.

Working with an AI agent

Extending the test setup and the analysis were all done with Claude Code and Opus 5.5 (high effort) within VS Code. My job was to set the direction, review the findings and make every decision. What Claude Opus 5.5 did:

  • Bisecting across releases. It pulled or rebuilt ten ch-elm images of the last nine months (Matchbox 4.0.16 to 4.1.18), ran the same load test on each, and found in which release validation became slower and in which it became faster again, to get a clearer view of the problem.

  • Profiling and heap dumps. With Java Flight Recorder it recorded which code paths parsed resources. With Eclipse Memory Analyzer it analyzed heap dumps in headless mode: which objects keep the memory alive, per package and per validation engine.

  • Reading upstream code, recommending and implementing optimizations. It compared Matchbox with the official HL7 validator in org.hl7.fhir.core, changed the code, ran the tests and repeated the measurements until the numbers and the validation results matched.

In 24 hours this added up to about 35 load and startup test runs, 17 Docker image builds, 10 Matchbox versions compared and 9 heap dumps analyzed.

The culprits in memory consumption

The heap dumps showed where the memory went. For ch-elm, 44,650 conformance resources (profiles, value sets, code systems) were loaded, all parsed into Java objects, consuming about 1 GB. The IG and its dependencies alone pull in seven versions of the HL7 terminology package. The main culprits and their fixes:

  1. Leftover package files (−171 MB). For a handful of small files, a helper kept a reference to the whole package it came from. So all raw files of hl7.fhir.r4.core and hl7.fhir.uv.xver-r5.r4 (210 MB) stayed in memory. It now keeps only the file it needs, which lowers the live heap by 171 MB.

  2. Identical strings (−15%, about 220 MB). The many package versions contain a lot of identical text. Java's string deduplication (-XX:+UseStringDeduplication) lets the garbage collector share it, at no measurable cost. It is now on by default.

  3. Everything parsed up front (−330 MB). The official validator loads resources lazily: it registers each resource with some metadata and parses it only when a validation needs it. Matchbox had never adopted this for its package loading. Now the terminology resources (code systems, value sets, naming systems, concept maps) are loaded lazily. Matchbox also pins the terminology resources of the FHIR core to their core versions; this step now reads the versions from the metadata instead of parsing every resource.

  4. Unused definitions of the database layer (−40 MB). The search indexing of the underlying HAPI FHIR server loaded all 649 FHIR core profiles for a type analysis it never uses.

The agent discovered all four of these issues and presented approaches to solve them. Not every idea worked. Loading all resources lazily made the first validation take 2.4 instead of 0.85 seconds, because the validator walks through all profiles on first use anyway. So profiles are parsed at startup, like in the official validator, and terminology on demand.

Results for a single IG

For ch-elm, a typical production setup with one IG, the ch-elm image 1.15.3 (based on Matchbox 4.1.18) needs 660 MB of heap. That is less than a third of image 1.13.1, which still ran with 3 GB, and a fifth of image 1.14.1, which needs 3.4 GB and no longer fits into 3 GB. It validates three times faster than 1.13.1, and as fast as 1.15.2. Compared with 1.15.2, the memory drops by more than half and the server starts about 25% faster. Startup is still slower than with 1.13.1; as each image changes both the Matchbox and the ch-elm version, the two effects can't be separated here.

Matchbox release (ch-elm image)4.0.16 (1.13.1)4.1.9 (1.14.1)4.1.17 (1.15.2)4.1.18 (1.15.3)
Heap limit of the image3 GB12 GB12 GB70% of the container memory
Live heap after 8,000 validations2,300 MB3,420 MB1,560 MB660 MB
Server ready after36–37 s43–46 s65–67 s48–49 s
Validation under load (median of 8,000)327 ms250 ms104 ms108 ms
Runs the load test in a 3 GB heapYesNo (live heap 3.4 GB)YesYes, also in 1 GB

Several IGs: share what engines have in common

Matchbox creates one validation engine per IG/validation parameters, so that the rules of one IG cannot interfere with another. The heap dumps showed that, apart from the FHIR core and the base terminology, each engine loaded all its packages again. Swiss IGs share most of their dependencies, so with twelve IGs, 133,000 resource objects existed more than once: 2.3 GB of duplicates. Now a cache shares the loaded packages between the engines. It keeps only weak references, so a package is freed when the last engine using it is dropped after an hour without use.

We measured this with with-preload, a sample configuration of Matchbox that preloads 12 Swiss IGs with their dependencies (55 packages). With the new default in a container with a 4 GB limit (a heap of 2.8 GB), all 12 engines ran 400 validations without failures and used 3.5 GB of the container's memory.

with-preload: 12 Swiss IGs, 55 packagesRelease 4.1.9Release 4.1.17Release 4.1.18
Heap limit of the image12 GB12 GB70% of the container memory
Live heap with all 12 engines10,630 MB4,756 MB1,187 MB
Memory held by the 12 IG engines alone8,123 MB3,465 MB177 MB
Duplicated between enginesn.a.2,278 MB1 MB
Creating an IG engine (median)19 s48 s16 s

What changes for you as a Matchbox user

The changes are released in Matchbox 4.1.18 and the ch-elm image 1.15.3:

  • A new memory default. The Docker image no longer sets a fixed 12 GB heap, which could even exceed the memory limit of a container. It now uses 70% of the container's memory limit (-XX:MaxRAMPercentage=70).

  • 4 GB is enough for a typical setup. We recommend a memory limit of 4 GB for the container, also for multiple IGs. A single IG such as ch-elm runs comfortably with 3 GB.

  • Configurable JVM options. The options can be overridden with the JDK_JAVA_OPTIONS environment variable (#594). Setting it replaces the default, so include a heap setting, -XX:+ExitOnOutOfMemoryError and -XX:+UseStringDeduplication.

  • A restart instead of a zombie. On the first OutOfMemoryError, the server now exits (-XX:+ExitOnOutOfMemoryError) instead of continuing in an undefined state while the health check reports it as up, so that Docker or Kubernetes can restart it (#457).

  • Faster engines. Compared with 4.1.17, the server starts about 25% faster and creates IG engines three times faster.

docker run -d --name matchbox -p 8080:8080 -m 4g europe-west6-docker.pkg.dev/ahdis-ch/ahdis/matchbox:latest

Lessons: a day instead of weeks with an agent

I could have done almost all of this myself. I'm familiar with load testing, bisecting releases, profiling and heap dump analysis, but less so with tools such as JFR or Eclipse Memory Analyzer. Even after we noticed the memory problem, I could not find the time: each step means waiting for a 30-minute test run, writing heap dump queries or reading code in a large upstream project.

Opus 5.5 took over exactly these steps: it ran them continuously, one after another and in parallel, in one session of 1 million tokens. I let it write a whole runbook and continued in a new session the next day, to verify that the setup can be reused and this is not a one-shot optimization.

The agent found features in the core libraries that we had missed, and proposed new approaches, such as sharing the loaded FHIR packages between the engines. I had not seen these qualities in agentic coding before. Opus 4.7 was the first model that could handle the Matchbox code, but Opus 5.5 is another leap: it can be used for software engineering, not just coding.

The reproducible test setup made this possible. Every proposal of the agent was measured against the same load test on the same machine, and every run checked that the validation results stayed the same. Every pull request goes through our Matchbox engine and server tests, which check specific validation features, and our integration tests, which check that the examples of the IGs we are interested in are still validated correctly. Without that, we could neither have trusted the improvements nor rejected the wrong ideas.

And a human in the loop stays essential. The agent proposed, measured and implemented, but not every proposal was right. Loading only the newest version of each terminology package would have saved 400 MB, but the IGs depend on specific versions, and validating against just the latest terminology version can change the results. Loading all resources lazily saved memory but made the first validation three times slower. I knew that multiple package versions are needed, which trade-offs are acceptable, and when a proposal was wrong.

In one day we delivered a Matchbox release that needs a fraction of the memory. But it needs the community to raise awareness and bring urgency to the problems, so thanks again to Achraf (@achrafachkari), Valentin (@reva) and especially my colleague Quentin (@qligier), co-maintainer of Matchbox, who has to review my pull requests co-authored with Claude Code.

Disclaimer: As with the code, I used Claude to help me with this blog article, but I read and refined every sentence.

Weiter
Weiter

Digital Health Terminology service in Switzerland