Java
    September 29, 202625 min read

    Java 25's Compact Object Headers: I Measured the Real Savings So You Don't Have To

    JEP 519 shrinks every Java object's header from 12 bytes to 8 in JDK 25. I ran real benchmarks on JDK 25.0.3 (not borrowed numbers) to measure the actual per-object savings for a realistic Spring Boot DTO shape, and explain why my numbers don't match the JEP's own headline figure.

    Share
    NOTE

    Compact Object Headers (JEP 519) shipped as a product feature in JDK 25 — no --enable-preview, no -XX:+UnlockExperimentalVMOptions, just -XX:+UseCompactObjectHeaders. JEP 534 will make it the default in JDK 27. If you're running JDK 25 or 26 today, this is one flag away.

    Every object you allocate in Java carries a header before a single field is stored — a mark word (locking/hashcode/GC state) plus a compressed class pointer. That header has been 12 bytes since compressed oops arrived in Java 6. Every JPA entity, every DTO, every element in a List<Order> pays it, whether the object needs 4 bytes of actual data or 400.

    JEP 519 shrinks it to 8 bytes. That sounds small. At the scale a real Spring Boot service runs at — millions of entities and DTOs live in the heap at any given moment — it isn't.

    I didn't want to just repeat the JEP's own headline number. So I ran real benchmarks on the JDK actually installed on this machine (Temurin JDK 25.0.3 LTS) and measured what this specific class shape actually saves. The numbers below are the literal output of that run — not adjusted, not rounded to look nicer.


    What Actually Changed

    An ordinary Java object's layout, with compressed class pointers (the default on heaps under ~32GB):

    graph TD subgraph "Before — 12-byte header" A1[Mark Word — 8 bytes] --> A2[Compressed Klass Pointer — 4 bytes] A2 --> A3[Fields...] end subgraph "After JEP 519 — 8-byte header" B1[Mark Word — compressed — 4 bytes] --> B2[Compressed Klass Pointer — 4 bytes] B2 --> B3[Fields...] end

    JEP 519 compresses the mark word itself — the same trick that compressed class pointers already used — folding lock state, identity hashcode, and GC age bits into a smaller footprint. This isn't free: some access patterns (biased locking edge cases, certain hash-code operations) pay a small CPU cost to decompress the mark word. The JEP is explicit about this tradeoff — it's a memory-for-CPU trade, not a strictly-better-in-every-dimension change. For the overwhelmingly common case (allocate, read fields, garbage collect), the memory win dominates.


    The Benchmark — Full Source, Actually Compiled and Run

    No abbreviated snippets here — this is the complete, unmodified source I compiled and ran, exactly as shown. Copy it, save it, run it yourself.

    NOTE

    Tested on: Eclipse Temurin JDK 25.0.3+9 LTS (openjdk version "25.0.3" 2026-04-21 LTS), Windows, single-file source launch (java Foo.java, no build tool). Every run below was executed twice to confirm the numbers were stable, not noise — both runs are shown.

    // ObjectHeaderBenchmark.java — full source, no external dependencies
    import java.util.ArrayList;
    import java.util.List;
    
    public class ObjectHeaderBenchmark {
    
        // Shape: long id, int quantity, double price, boolean active,
        // long timestampEpoch, String sku, String status — an order-line DTO.
        static final class OrderLine {
            final long id;
            final int quantity;
            final double price;
            final boolean active;
            final long timestampEpoch;
            final String sku;
            final String status;
    
            OrderLine(long id, int quantity, double price, boolean active, long timestampEpoch, String sku, String status) {
                this.id = id;
                this.quantity = quantity;
                this.price = price;
                this.active = active;
                this.timestampEpoch = timestampEpoch;
                this.sku = sku;
                this.status = status;
            }
        }
    
        static long usedHeapBytes() {
            Runtime rt = Runtime.getRuntime();
            return rt.totalMemory() - rt.freeMemory();
        }
    
        static void settle() throws InterruptedException {
            for (int i = 0; i < 3; i++) {
                System.gc();
                Thread.sleep(150);
            }
        }
    
        public static void main(String[] args) throws Exception {
            int n = args.length > 0 ? Integer.parseInt(args[0]) : 8_000_000;
    
            System.out.println("JVM: " + System.getProperty("java.vm.name") + " " + System.getProperty("java.version"));
    
            settle();
            long before = usedHeapBytes();
    
            List<OrderLine> lines = new ArrayList<>(n);
            for (int i = 0; i < n; i++) {
                // Per-instance-distinct strings via concatenation — guarantees a
                // real, non-interned heap allocation for every instance, not a
                // shared constant-pool object that would understate the effect.
                String sku = "SKU-" + i;
                String status = "STATUS-" + (i % 7);
                lines.add(new OrderLine(i, (i % 50) + 1, 9.99 + (i % 500), (i % 3 == 0), 1_700_000_000L + i, sku, status));
            }
    
            settle();
            long after = usedHeapBytes();
            long delta = after - before;
    
            System.out.println("N (OrderLine instances)        = " + n);
            System.out.println("Heap used before allocation     = " + before + " bytes");
            System.out.println("Heap used after allocation+GC   = " + after + " bytes");
            System.out.println("Delta                           = " + delta + " bytes");
            System.out.println("Bytes per OrderLine instance    = " + String.format("%.2f", (double) delta / n));
    
            if (lines.isEmpty()) throw new AssertionError("unreachable, keeps 'lines' live");
        }
    }

    Run exactly like this — nothing hidden:

    java -Xms4g -Xmx4g -XX:-UseCompactObjectHeaders ObjectHeaderBenchmark.java 8000000
    java -Xms4g -Xmx4g -XX:+UseCompactObjectHeaders ObjectHeaderBenchmark.java 8000000

    Actual console output, run 1 — default 12-byte header:

    JVM: OpenJDK 64-Bit Server VM 25.0.3
    N (OrderLine instances)        = 8000000
    Heap used before allocation     = 4642776 bytes
    Heap used after allocation+GC   = 1318189424 bytes
    Delta                           = 1313546648 bytes
    Bytes per OrderLine instance    = 164.19

    Actual console output, run 1 — -XX:+UseCompactObjectHeaders:

    JVM: OpenJDK 64-Bit Server VM 25.0.3
    N (OrderLine instances)        = 8000000
    Heap used before allocation     = 4350224 bytes
    Heap used after allocation+GC   = 1190352392 bytes
    Delta                           = 1186002168 bytes
    Bytes per OrderLine instance    = 148.25

    Repeat run, both configurations, to confirm the numbers weren't a fluke: 164.20 bytes/instance (default) and 148.21 bytes/instance (compact) — stable to within 0.02%.

    Configuration Bytes per OrderLine instance (3 heap objects: the DTO + its 2 owned Strings)
    Default (12-byte header) 164.19 / 164.20
    -XX:+UseCompactObjectHeaders 148.25 / 148.21
    Measured savings ≈15.95 bytes per instance

    I also ran a second, deliberately simpler variant — a primitives-only class with no nested objects, so exactly one heap object is allocated per instance, isolating the header effect from any nested-string savings:

    // ObjectHeaderBenchmarkSimple.java — full source, companion isolation test
    import java.util.ArrayList;
    import java.util.List;
    
    public class ObjectHeaderBenchmarkSimple {
    
        static final class Metric {
            final long id;
            final int count;
            final double value;
            final boolean active;
            final long timestampEpoch;
    
            Metric(long id, int count, double value, boolean active, long timestampEpoch) {
                this.id = id;
                this.count = count;
                this.value = value;
                this.active = active;
                this.timestampEpoch = timestampEpoch;
            }
        }
    
        static long usedHeapBytes() {
            Runtime rt = Runtime.getRuntime();
            return rt.totalMemory() - rt.freeMemory();
        }
    
        static void settle() throws InterruptedException {
            for (int i = 0; i < 3; i++) { System.gc(); Thread.sleep(150); }
        }
    
        public static void main(String[] args) throws Exception {
            int n = args.length > 0 ? Integer.parseInt(args[0]) : 8_000_000;
    
            settle();
            long before = usedHeapBytes();
    
            List<Metric> items = new ArrayList<>(n);
            for (int i = 0; i < n; i++) {
                items.add(new Metric(i, i % 50, 9.99 + (i % 500), (i % 3 == 0), 1_700_000_000L + i));
            }
    
            settle();
            long after = usedHeapBytes();
            long delta = after - before;
    
            System.out.println("N (Metric instances, single heap object each) = " + n);
            System.out.println("Delta                                          = " + delta + " bytes");
            System.out.println("Bytes per instance                             = " + String.format("%.2f", (double) delta / n));
    
            if (items.isEmpty()) throw new AssertionError("unreachable, keeps 'items' live");
        }
    }

    Actual console output — default:

    N (Metric instances, single heap object each) = 8000000
    Delta                                          = 417675096 bytes
    Bytes per instance                             = 52.21

    Actual console output — -XX:+UseCompactObjectHeaders:

    N (Metric instances, single heap object each) = 8000000
    Delta                                          = 353674232 bytes
    Bytes per instance                             = 44.21
    Configuration Bytes per Metric instance (1 heap object)
    Default (12-byte header) 52.21
    -XX:+UseCompactObjectHeaders 44.21
    Measured savings exactly 8.00 bytes per instance

    Why My Numbers Don't Match "4 Bytes Per Object"

    The JEP's own framing is "reduces the Java heap footprint by 4 bytes per object on average." My clean, single-heap-object benchmark measured 8 bytes saved, not 4 — double the headline figure. That's not a benchmarking error; it's alignment.

    The JVM rounds every object's total size up to an 8-byte boundary. Shrinking the header by 4 bytes doesn't always shrink the rounded object size by exactly 4 bytes — it depends on where the unpadded size falls relative to that boundary. My Metric class's field layout happens to land such that the smaller header pushes the whole object across an alignment boundary, saving a full 8-byte slot instead of a partial 4. A class with different field ordering or count could just as easily save 0 bytes if its padding already absorbed the difference, or exactly 4 if it lands differently.

    This is the actual lesson: "4 bytes per object" is a fleet-wide average across a real, mixed codebase — your specific classes will save 0, 4, or 8 bytes depending on their field layout and alignment, not a uniform 4. Don't take either number — mine or the JEP's — as what your classes will show. Measure your own shapes the way I measured mine; the benchmark above is under 70 lines and runs in seconds.

    The OrderLine result (≈15.95 bytes across 3 owned heap objects, ≈5.3 bytes/object average) sits between the two single-class numbers precisely because it's a blend of the DTO's own layout and its two Strings' layout, each hitting the alignment boundary differently.


    What This Means at Spring Boot Scale

    Scope this honestly: it shrinks object header overhead specifically. It is not a "your app uses 30% less memory" claim, and anyone telling you that from this JEP alone is overselling it.

    But translated to a realistic number: a service holding 20 million cached order-line objects (a plausible in-memory cache or a large batch-processing window) at the measured ≈15.95 bytes/instance saving:

    20,000,000 × 15.95 bytes ≈ 319,000,000 bytes ≈ 304 MiB saved

    That's real heap headroom recovered from a one-line JVM flag change, with zero code changes, zero redeploy risk beyond the flag itself. For a heap-constrained service running close to its container memory limit, 300+ MiB is the difference between comfortable and paging.

    The effect compounds with anything that holds large object graphs: JPA persistence contexts with big result sets, in-memory caches (Caffeine, Guava), large List<T>/Map<K,V> collections, and — relevant to this site specifically — the document chunks and embeddings held in memory by a RAG pipeline built with Spring AI.


    Pros and Cons

    Pros Cons
    Real, measured memory savings (8–16 bytes/instance in my benchmarks) with zero code changes Savings are alignment-dependent — some class shapes save 0 bytes, not a guaranteed uniform win
    One JVM flag, instantly reversible — no redeploy risk beyond a restart Small CPU cost decompressing the mark word on some access patterns (biased locking edge cases, certain hashCode() paths)
    Compounds automatically with anything holding large object graphs — caches, JPA result sets, RAG embeddings Breaks code that reads raw object header bytes (JNI, some Unsafe-based tooling) — rare, but a real compatibility risk
    Becomes the JDK 27 default anyway (JEP 534) — adopting now means you're ahead of a change that's coming regardless Benefit shrinks as a % of total size for already-large objects — a 4–8 byte header shrink barely registers on a 10KB object
    No API surface change — nothing to migrate, no new imports, no dependency bump Less battle-tested than G1 or virtual threads — GA only since JDK 25, so production track record is still short

    When to Use It vs. When Not To

    Use it when:

    • Your workload is object-header-heavy — lots of small objects (DTOs, entities, value objects, boxed wrappers), not a few large ones. Header overhead is proportionally biggest exactly where you have the most instances.
    • You're memory-constrained — a container memory limit you're pushing against, or GC pause pressure from a large heap — and want a free lever to try before reaching for an architecture change.
    • You're already on JDK 25+ with no low-level JNI/Unsafe dependency that touches object layout directly.
    • You can canary it — you have the observability and rollback path to validate before a fleet-wide rollout (see below). This is what makes it low-risk to even experiment with.

    Don't enable this yet if:

    • You depend on native code (JNI) or a library that inspects raw object header bytes directly — anything hand-rolled that assumes the pre-519 12-byte layout can misbehave. The JEP calls this out explicitly as the primary compatibility risk.
    • You use a serialization library with header-layout assumptions baked in (rare, but some low-level off-heap/unsafe-based tools exist). Check your dependency list, don't assume.
    • Your objects are already large — a service dominated by big byte[] buffers or large arrays won't feel this; the header is a small fraction of an already-large object, so the win is diluted to near-zero.
    • You haven't measured your own workload yet. The alignment effect above means your savings could be 0 bytes for some classes — verify before you write "we cut memory by X%" in a postmortem.
    • You're on a JDK older than 25, or a vendor build that hasn't backported JEP 519 — check java -XX:+PrintFlagsFinal -version | grep CompactObjectHeaders before assuming it's available.

    How to try it safely: flip the flag on one canary instance behind your existing traffic split, watch GC pause times and heap-used metrics in your existing observability stack (the same OpenTelemetry setup you already have) for a few days, then roll forward. It's a JVM flag, not a code change — rollback is instant.


    A Real Use Case: The Nightly Reconciliation Job

    Here's where this actually mattered in practice, using the same OrderLine shape from the benchmark above.

    Picture a nightly batch job that reconciles the day's orders: it loads every order line from the last 24 hours into memory — say 12 million rows on a busy day — builds an in-memory index keyed by SKU, cross-checks it against a warehouse feed, and writes discrepancies back. The whole job runs inside one JVM, holding all 12 million OrderLine-shaped objects live at once because the cross-check needs random access across the full set, not a streaming pass.

    Before: at the measured 164.19 bytes/instance, 12 million rows costs ≈1.83 GiB just for this in-memory index — before the job's own working set, before the warehouse-feed data it's comparing against, before JVM/GC overhead on top.

    After enabling -XX:+UseCompactObjectHeaders: at 148.25 bytes/instance, the same 12 million rows costs ≈1.66 GiB — ≈182 MiB recovered, with the job's code completely unchanged.

    That 182 MiB was the difference, in this shape of scenario, between the job fitting comfortably inside its container's memory limit and occasionally tripping the OOM killer on the busiest reconciliation nights — the nights with the most orders, which are exactly the nights you can least afford the job to fail. The fix here isn't "rewrite the index to be more memory-efficient" (a real project, with real risk) — it's one flag on the job's JVM invocation, validated on a few non-critical runs first, then rolled into the production cron entry.

    This is the pattern worth internalizing: this JEP doesn't fix a bad architecture, but it gives you real, measurable headroom on an unchanged architecture — which is exactly the kind of low-risk win worth taking before reaching for a bigger redesign.


    Try It Yourself

    Both full source files are above, unmodified from what I actually ran — save either one, verbatim, into a file matching its class name (ObjectHeaderBenchmark.java / ObjectHeaderBenchmarkSimple.java), and run:

    java -Xms4g -Xmx4g -XX:-UseCompactObjectHeaders ObjectHeaderBenchmark.java 8000000
    java -Xms4g -Xmx4g -XX:+UseCompactObjectHeaders ObjectHeaderBenchmark.java 8000000

    No build tool, no dependencies — single-file source launch runs it directly. Then swap the OrderLine field list for your own entity/DTO shape and re-run. That's the only number worth putting in a capacity-planning doc — not mine, not the JEP's, yours.


    Resources


    Ran these benchmarks yourself and got different numbers? That's expected — send me your class shape and results. Find me on LinkedIn or browse the full buildingai.in blog.

    Ask about this article

    Get answers grounded in this post. AI-generated — based on this article, and may be imperfect.

    Was this helpful?
    AY
    Avaneesh Yadav

    I build enterprise AI systems — Spring AI, RAG, and agents — and write about shipping LLMs to production. I also run advisory and workshops for engineering teams.

    Scaled AI Weekly

    Enjoyed this? Get more like it every Monday.

    Real architecture decisions, LLMOps patterns that survive production, and engineering leadership advice — from 12+ years of building at enterprise scale. Free. No spam. Unsubscribe anytime.

    Join engineers building production AI systems

    Free: LLM Production Readiness Checklist (PDF)

    50 checks across observability, rate limiting, cost optimization, failure handling, and security — for teams shipping AI features to production.

    No spam. Unsubscribe any time.

    Comments