The .class File Format: Header, Constant Pool, Fields, and Methods
Objective
Understand the .class file: the binary format javac produces from Java source and the only thing the JVM actually loads — a fixed sequence of sections (magic number, version, constant pool, access flags, fields, methods, attributes) that lets the JVM verify and interpret a class without ever seeing the original .java code.
Use Cases
- Reading
javap -voutput to understand what a class actually compiled to, instead of guessing from the source. - Diagnosing
UnsupportedClassVersionErrorby knowing what the class file's major version number means and how it maps to a Java release. - Understanding why decompilers, bytecode-manipulation libraries (ASM, ByteBuddy), and frameworks that do classpath scanning all start by parsing the same fixed layout.
- Explaining why an object's field or method names show up in error messages and stack traces even though the JVM "doesn't understand Java" — they're stored as UTF-8 entries in the constant pool.
- Reading a raw descriptor like
[[Ljava/lang/String;or(II)Ioff a stack trace or disassembly without needingjavapto spell it out in plain English. - Explaining why
List<String>andList<Integer>are indistinguishable at the bytecode level — both compile to the same erased descriptor.
Deep Dive
From source to bytecode: the compiler pipeline
The JVM never reads .java source. javac compiles it into a .class file, and that binary file — not the original code — is what the class loader reads:
plaintextHelloWorld.java → javac → HelloWorld.class → JVM class loader → execution
Every .class file, regardless of what the source looked like, is laid out as the same fixed sequence of sections: magic number, version, constant pool, access flags, this class / super class, interfaces, fields, methods, attributes.
Magic Number: identifying a valid class file
The first 4 bytes of every .class file are a fixed signature, CAFEBABE, checked before anything else is parsed. It's visible in a raw hex dump, but not as a labeled Magic: line in javap -v output on current JDKs — verbose javap instead reports the file's on-disk metadata (last-modified time, size, and a SHA-256 checksum of the bytes), then moves straight into the class declaration:
plaintext$ xxd HelloWorld.class | head -1 00000000: cafe babe 0000 0041 0013 0a00 0200 0307 .......A........
plaintext$ javap -v HelloWorld.class | head -4 Classfile /home/user/HelloWorld.class Last modified Aug 19, 2026; size 428 bytes SHA-256 checksum 3a1f...e29c Compiled from "HelloWorld.java"
If those first 4 bytes don't match — a truncated download, a text file renamed to .class — the JVM throws ClassFormatError before attempting to read anything else in the file. javap's SHA-256 line (added via -sysinfo, which -v implies) is a convenience for confirming file integrity — it plays no role in class loading itself; only the verifier's own checks, starting with the magic number, decide whether the JVM accepts the file.
Version: minor and major
Right after the magic number come two 2-byte fields, minor_version and major_version. major_version identifies the bytecode format and increases with the Java release that introduced it:
| major | Java |
|---|---|
| 52 | Java 8 |
| 55 | Java 11 |
| 61 | Java 17 |
| 65 | Java 21 |
plaintext$ javap -v HelloWorld.class | grep version minor version: 0 major version: 65
The JVM compares this number against what it supports at load time, before running a single instruction from the file.
Constant Pool: the class's table of symbolic references
The constant pool is a table of every class name, method signature, field name, string literal, and numeric constant the class refers to. Nothing else in the file stores those values directly — they're all referenced by index into this table:
javapublic class HelloWorld {
public static void main(String[] args) {
System.out.println("Hello");
}
}plaintext$ javap -v HelloWorld.class Constant pool: #1 = Methodref #6.#15 // java/lang/Object."<init>":()V #2 = Fieldref #16.#17 // java/lang/System.out:Ljava/io/PrintStream; #3 = String #18 // Hello #4 = Methodref #19.#20 // java/io/PrintStream.println:(Ljava/lang/String;)V #5 = Class #21 // HelloWorld ...
The bytecode for main doesn't contain the string "Hello" or the class name java.io.PrintStream inline — it references constant pool entries #3 and #2 by index. The pool is a catalog other sections of the file point into, not the program itself.
Access Flags
access_flags is a bitmask right after the constant pool describing the class itself:
plaintext$ javap -v HelloWorld.class | grep flags flags: (0x0021) ACC_PUBLIC, ACC_SUPER
| flag | meaning |
|---|---|
ACC_PUBLIC |
the class is public |
ACC_FINAL |
the class cannot be extended |
ACC_SUPER |
historical flag affecting invokespecial resolution for superclass method calls |
ACC_INTERFACE |
the file describes an interface, not a class |
ACC_ABSTRACT |
the class is abstract |
ACC_SYNTHETIC |
the class was generated by the compiler, not written directly in source |
Fields: instance vs. static
A field's entry (field_info) stores a name index and a descriptor index into the constant pool, plus its own access_flags. Comparing an instance field to a static one shows the only structural difference is that flag:
javapublic class Counter {
int value; // instance field
static int instances; // class field
}plaintext$ javap -p -v Counter.class | grep -A2 'value\|instances' int value; descriptor: I flags: (0x0000) static int instances; descriptor: I flags: (0x0008) ACC_STATIC
An instance field gets its own storage per object — each Counter has its own value. A static field is stored once on the class itself and shared by every instance — that's exactly what ACC_STATIC tells the JVM to do differently when allocating and resolving it.
Methods: parameters, return type, and bytecode
A method's entry stores its name, descriptor (parameter types + return type, encoded as a string), access flags, and — for anything with a body — a Code attribute holding the actual bytecode instructions:
javapublic int add(int a, int b) {
return a + b;
}plaintext$ javap -v Calc.class | grep -A6 'public int add' public int add(int, int); descriptor: (II)I flags: (0x0001) ACC_PUBLIC Code: stack=2, locals=3, args_size=3 0: iload_1 1: iload_2 2: iadd 3: ireturn
The descriptor (II)I says "takes two ints, returns an int" — V in that position means void, and reference types use the fully-qualified Lpackage/Class; form, as seen earlier in (Ljava/lang/String;)V for println. The Code attribute is what the JVM actually executes; everything else in the file exists to let the JVM resolve and verify it correctly.
Field and method descriptors: the type-code alphabet
Every field and method descriptor in the constant pool is built from a small fixed set of one-letter codes for primitives, plus two structural prefixes for everything that isn't a primitive:
| code | type |
|---|---|
B |
byte |
C |
char |
D |
double |
F |
float |
I |
int |
J |
long |
S |
short |
Z |
boolean |
L ClassName ; |
a reference type — fully-qualified, slash-separated, terminated by ; |
[ |
one array dimension — prefixed onto whatever descriptor the element type has |
Compiling a sampler class and reading its field descriptors shows the pattern directly — array types just stack [ in front of the element descriptor, once per dimension:
javabyte b;
int[] intArray;
String[][] stringMatrix;plaintextbyte b; descriptor: B int[] intArray; descriptor: [I java.lang.String[][] stringMatrix; descriptor: [[Ljava/lang/String;
Generics don't get their own descriptor syntax at all — List<String> and List<Integer> both compile to the identical raw descriptor Ljava/util/List;. The generic type argument is preserved separately, in an optional Signature attribute (Ljava/util/List<Ljava/lang/String;>;) that only tools like javac and reflection consult; the bytecode itself, and the verifier, only ever see the erased Ljava/util/List;. This is type erasure made concrete at the descriptor level: the JVM has no instruction or descriptor code that distinguishes a List<String> from a List<Integer>.
Trade-offs
- Indirection vs. size — every symbolic reference in the bytecode is a constant-pool index rather than an inlined value, which lets the same string or method reference be reused across many instructions instead of duplicated, at the cost of a lookup at link/resolution time.
plaintext2: invokevirtual #4 // Method println:(Ljava/lang/String;)V — resolved through the pool, not inlined
- The version check is one-directional — a JVM refuses to load a class file whose
major_versionis newer than it supports, but happily loads a file compiled for an older release:
plaintext$ java HelloWorld Error: HelloWorld has been compiled by a more recent version of the Java Runtime (class file version 65.0), this version of the Java Runtime only recognizes class file versions up to 61.0
ACC_SUPERis a historical compatibility bit — every class compiled since Java 1.0.2 has it set automatically, existing only so a modern JVM can still correctly resolveinvokespecialcalls the way pre-1.0.2 class files expected; there's no reason to reason about it in code written today.- Constant pool entries are 1-indexed and never entry
#0— index0is reserved as an explicit "no reference" value (used, for example, by a class with no superclass), so pool entries always start counting at#1, which trips up anyone writing a parser by hand and assuming 0-based indexing.