← JVM Concepts

The .class File Format: Header, Constant Pool, Fields, and Methods

Published on 2026-08-10·v1.2
🔬 Open Lab

Objective

Understand the .class file: the binary format javac produces from Java source and the only thing the JVM actually loads — a fixed sequence of sections (magic number, version, constant pool, access flags, fields, methods, attributes) that lets the JVM verify and interpret a class without ever seeing the original .java code.

Use Cases

  • Reading javap -v output to understand what a class actually compiled to, instead of guessing from the source.
  • Diagnosing UnsupportedClassVersionError by knowing what the class file's major version number means and how it maps to a Java release.
  • Understanding why decompilers, bytecode-manipulation libraries (ASM, ByteBuddy), and frameworks that do classpath scanning all start by parsing the same fixed layout.
  • Explaining why an object's field or method names show up in error messages and stack traces even though the JVM "doesn't understand Java" — they're stored as UTF-8 entries in the constant pool.
  • Reading a raw descriptor like [[Ljava/lang/String; or (II)I off a stack trace or disassembly without needing javap to spell it out in plain English.
  • Explaining why List<String> and List<Integer> are indistinguishable at the bytecode level — both compile to the same erased descriptor.

Deep Dive

From source to bytecode: the compiler pipeline

The JVM never reads .java source. javac compiles it into a .class file, and that binary file — not the original code — is what the class loader reads:

plaintext
HelloWorld.java → javac → HelloWorld.class → JVM class loader → execution

Every .class file, regardless of what the source looked like, is laid out as the same fixed sequence of sections: magic number, version, constant pool, access flags, this class / super class, interfaces, fields, methods, attributes.

Magic Number: identifying a valid class file

The first 4 bytes of every .class file are a fixed signature, CAFEBABE, checked before anything else is parsed. It's visible in a raw hex dump, but not as a labeled Magic: line in javap -v output on current JDKs — verbose javap instead reports the file's on-disk metadata (last-modified time, size, and a SHA-256 checksum of the bytes), then moves straight into the class declaration:

plaintext
$ xxd HelloWorld.class | head -1 00000000: cafe babe 0000 0041 0013 0a00 0200 0307 .......A........
plaintext
$ javap -v HelloWorld.class | head -4 Classfile /home/user/HelloWorld.class Last modified Aug 19, 2026; size 428 bytes SHA-256 checksum 3a1f...e29c Compiled from "HelloWorld.java"

If those first 4 bytes don't match — a truncated download, a text file renamed to .class — the JVM throws ClassFormatError before attempting to read anything else in the file. javap's SHA-256 line (added via -sysinfo, which -v implies) is a convenience for confirming file integrity — it plays no role in class loading itself; only the verifier's own checks, starting with the magic number, decide whether the JVM accepts the file.

Version: minor and major

Right after the magic number come two 2-byte fields, minor_version and major_version. major_version identifies the bytecode format and increases with the Java release that introduced it:

major Java
52 Java 8
55 Java 11
61 Java 17
65 Java 21
plaintext
$ javap -v HelloWorld.class | grep version minor version: 0 major version: 65

The JVM compares this number against what it supports at load time, before running a single instruction from the file.

Constant Pool: the class's table of symbolic references

The constant pool is a table of every class name, method signature, field name, string literal, and numeric constant the class refers to. Nothing else in the file stores those values directly — they're all referenced by index into this table:

java
public class HelloWorld { public static void main(String[] args) { System.out.println("Hello"); } }
plaintext
$ javap -v HelloWorld.class Constant pool: #1 = Methodref #6.#15 // java/lang/Object."<init>":()V #2 = Fieldref #16.#17 // java/lang/System.out:Ljava/io/PrintStream; #3 = String #18 // Hello #4 = Methodref #19.#20 // java/io/PrintStream.println:(Ljava/lang/String;)V #5 = Class #21 // HelloWorld ...

The bytecode for main doesn't contain the string "Hello" or the class name java.io.PrintStream inline — it references constant pool entries #3 and #2 by index. The pool is a catalog other sections of the file point into, not the program itself.

Access Flags

access_flags is a bitmask right after the constant pool describing the class itself:

plaintext
$ javap -v HelloWorld.class | grep flags flags: (0x0021) ACC_PUBLIC, ACC_SUPER
flag meaning
ACC_PUBLIC the class is public
ACC_FINAL the class cannot be extended
ACC_SUPER historical flag affecting invokespecial resolution for superclass method calls
ACC_INTERFACE the file describes an interface, not a class
ACC_ABSTRACT the class is abstract
ACC_SYNTHETIC the class was generated by the compiler, not written directly in source

Fields: instance vs. static

A field's entry (field_info) stores a name index and a descriptor index into the constant pool, plus its own access_flags. Comparing an instance field to a static one shows the only structural difference is that flag:

java
public class Counter { int value; // instance field static int instances; // class field }
plaintext
$ javap -p -v Counter.class | grep -A2 'value\|instances' int value; descriptor: I flags: (0x0000) static int instances; descriptor: I flags: (0x0008) ACC_STATIC

An instance field gets its own storage per object — each Counter has its own value. A static field is stored once on the class itself and shared by every instance — that's exactly what ACC_STATIC tells the JVM to do differently when allocating and resolving it.

Methods: parameters, return type, and bytecode

A method's entry stores its name, descriptor (parameter types + return type, encoded as a string), access flags, and — for anything with a body — a Code attribute holding the actual bytecode instructions:

java
public int add(int a, int b) { return a + b; }
plaintext
$ javap -v Calc.class | grep -A6 'public int add' public int add(int, int); descriptor: (II)I flags: (0x0001) ACC_PUBLIC Code: stack=2, locals=3, args_size=3 0: iload_1 1: iload_2 2: iadd 3: ireturn

The descriptor (II)I says "takes two ints, returns an int" — V in that position means void, and reference types use the fully-qualified Lpackage/Class; form, as seen earlier in (Ljava/lang/String;)V for println. The Code attribute is what the JVM actually executes; everything else in the file exists to let the JVM resolve and verify it correctly.

Field and method descriptors: the type-code alphabet

Every field and method descriptor in the constant pool is built from a small fixed set of one-letter codes for primitives, plus two structural prefixes for everything that isn't a primitive:

code type
B byte
C char
D double
F float
I int
J long
S short
Z boolean
L ClassName ; a reference type — fully-qualified, slash-separated, terminated by ;
[ one array dimension — prefixed onto whatever descriptor the element type has

Compiling a sampler class and reading its field descriptors shows the pattern directly — array types just stack [ in front of the element descriptor, once per dimension:

java
byte b; int[] intArray; String[][] stringMatrix;
plaintext
byte b; descriptor: B int[] intArray; descriptor: [I java.lang.String[][] stringMatrix; descriptor: [[Ljava/lang/String;

Generics don't get their own descriptor syntax at all — List<String> and List<Integer> both compile to the identical raw descriptor Ljava/util/List;. The generic type argument is preserved separately, in an optional Signature attribute (Ljava/util/List<Ljava/lang/String;>;) that only tools like javac and reflection consult; the bytecode itself, and the verifier, only ever see the erased Ljava/util/List;. This is type erasure made concrete at the descriptor level: the JVM has no instruction or descriptor code that distinguishes a List<String> from a List<Integer>.

Trade-offs

  • Indirection vs. size — every symbolic reference in the bytecode is a constant-pool index rather than an inlined value, which lets the same string or method reference be reused across many instructions instead of duplicated, at the cost of a lookup at link/resolution time.
plaintext
2: invokevirtual #4 // Method println:(Ljava/lang/String;)V — resolved through the pool, not inlined
  • The version check is one-directional — a JVM refuses to load a class file whose major_version is newer than it supports, but happily loads a file compiled for an older release:
plaintext
$ java HelloWorld Error: HelloWorld has been compiled by a more recent version of the Java Runtime (class file version 65.0), this version of the Java Runtime only recognizes class file versions up to 61.0
  • ACC_SUPER is a historical compatibility bit — every class compiled since Java 1.0.2 has it set automatically, existing only so a modern JVM can still correctly resolve invokespecial calls the way pre-1.0.2 class files expected; there's no reason to reason about it in code written today.
  • Constant pool entries are 1-indexed and never entry #0 — index 0 is reserved as an explicit "no reference" value (used, for example, by a class with no superclass), so pool entries always start counting at #1, which trips up anyone writing a parser by hand and assuming 0-based indexing.

Documentation Links