MongoDB Monitoring and Backup Strategies
Objective
Understand what actually needs watching on a running MongoDB deployment before it becomes an incident, and how to get data out safely when something goes wrong anyway. The book frames monitoring as a prerequisite, not an afterthought: "Before you deploy, it is important to set up some type of monitoring. Monitoring should allow you to track what your server is doing and alert you if something goes wrong." Backups get the same framing on the other side of the same coin — "Backups are good protection against most types of failure, and very little can't be solved by restoring from a clean backup," but only if you've actually practiced the restore, not just the dump.
Use Cases
- Running
mongostatduring a live incident to get a one-line-per-second view of ops/sec, page faults, queue depth, and connections across the whole deployment, without opening a dashboard. - Running
mongotopto find which specific collection is consuming read/write time right now, whenmongostatsays the server is busy but doesn't say with what. - Reading
db.serverStatus()'sextra_info.page_faultsandwiredTiger.concurrentTransactionsfields to get the same numbers a dashboard would show, scriptable and without any third-party agent installed. - Watching replication lag graphs to catch a secondary silently falling behind — the standard warning sign of an overloaded system, a missing
_idindex, or a network problem — before it goes stale and needs a full resync. - Choosing a backup method by deployment shape:
mongodump/mongorestorefor a single database or collection, filesystem/volume snapshots for a whole replica set member quickly and consistently, and a cloud provider's managed backup service (Atlas, or Cloud/Ops Manager for self-hosted) when the whole cluster — including a sharded cluster's config servers — needs to be covered on a schedule.
Deep Dive
The tool ladder: mongostat, mongotop, serverStatus, and a real monitoring system
The book lays out monitoring as three questions worth tracking continuously: "How to track MongoDB's memory usage. How to track application performance metrics. How to diagnose replication issues." It demonstrates the metrics using MongoDB Ops Manager graphs, noting "the monitoring capabilities of MongoDB Atlas (MongoDB's cloud database service) are very similar," and it names a free option too: MongoDB's free monitoring service, which "keeps the monitoring data for 24 hours after it has been uploaded and provides coarse-grained statistics on operation execution times, memory usage, CPU usage, and operation counts." The chapter's closing advice is blunt: "If you do not want to use Ops Manager, Atlas, or MongoDB's free monitoring service, please use some type of monitoring."
Underneath any dashboard, the same three primitives supply the numbers: mongostat (a vmstat-style per-second summary of the whole server — ops/sec by type, page faults, queued reads/writes, connections, replication lag), mongotop (the same idea narrowed to per-collection read/write time, useful for finding which collection is hot), and db.serverStatus() (a single JSON snapshot of everything, including fields the book quotes directly — db.adminCommand({"serverStatus": 1})["extra_info"] returning { "note" : "fields vary by platform", "page_faults" : 50 }, a running count of page faults since the process started).
Memory and page faults: the metric that matters most
The book's reasoning here is foundational to everything else in the chapter: "Accessing data in memory is fast, and accessing data on disk is slow... typically MongoDB uses up memory before any other resource... MongoDB's memory usage is one of the most important stats to track." It breaks server memory into three reported values — resident (data actually paged into RAM right now), virtual (the OS-level address space abstraction, "typically twice the size of the mapped memory"), and mapped (relevant only to the pre-4.0 MMAP storage engine — "now that MongoDB uses the WiredTiger storage engine, you should see zero usage for mapped memory"). Resident memory is the one that answers the real question: "If your data fits entirely in memory, the resident memory should be approximately the size of your data."
Page faults are the leading indicator of memory pressure: a page fault means "the data MongoDB is looking for is not in RAM" and has to be copied from disk, which is orders of magnitude slower. The book's guidance on interpreting the number is deliberately non-absolute — "if the disk... can handle that many faults and the application can handle the delay of the disk seeks, there is no particular problem with having so many faults" — but it flags disk load as nonlinear: "once a disk begins getting overloaded, each operation must queue for a longer and longer period of time, creating a chain reaction." The actionable version of this: "Track your page fault numbers over time. If your application is behaving well with a certain number of page faults, you have a baseline for how many page faults the system can handle. If page faults begin to creep up and performance deteriorates, you have a threshold to alert on." In other words, page faults aren't a metric with a universal red line — they're a metric you baseline for your workload and alert on deviation from.
The concept underneath all of it is the working set — "the data and indexes that your application uses," which the book ranks by desirability: entire dataset in memory (best, often infeasible), working set in memory (the common target — "generally there's a core dataset... that covers 90% of requests"), down to no useful subset in memory ("it will be slow"). Sizing a working set is a measurement exercise, not a guess: track how much data is created and how much of it stays "hot" — the book's worked example puts a 2 GB/week write rate with a month-long access pattern at roughly a 3.2 GB working set, "plus a fudge factor for indexes."
Replication lag and oplog length
Once memory is understood, replication health is the next tier. Lag is defined precisely: "It's calculated by subtracting the time of the last op applied on a secondary from the time of the last op on the primary." A healthy secondary shows lag "basically 0 all the time," on the order of milliseconds. The two failure shapes the book describes are worth distinguishing because they call for different responses:
- Stuck replication — lag grows by exactly one second per second, "the most extreme case," typically caused by a network problem or a missing
_idindex (which the book calls out as required for replication to function). The fix it prescribes is specific: take the secondary out of the set, start it standalone, and build the index — "as a unique index" — before rejoining. - Gradual overload — a secondary slowly falls behind under sustained load without the telltale one-second-per-second slope, because "some replication will still be happening." This is a capacity problem, not a corruption or config problem, and the fix is reducing load on that member, not repairing it.
The book also flags a read: "sudden spikes in replication lag" on a very low-write system can be phantom lag — a sampling artifact from measuring the secondary's timestamp right before the next primary write, not real lag. Don't chase a spike that disappears once the write rate goes up.
Oplog length — how much wall-clock time a member's oplog covers before it starts overwriting old entries — is the safety margin for how long a secondary can be down or lagging before it goes stale and needs a full resync. The book's rule of thumb: "Every member that might become primary should have an oplog longer than a day... In general, oplogs should be as long as you can afford the disk space to make them" — the trade is genuinely one-sided, since "they take up basically no memory."
Lock percentage's successor: the ticketing system
Older MongoDB monitoring dashboards tracked a "lock percentage" metric under the MMAP storage engine's global lock model. WiredTiger replaced that with document-level concurrency plus a ticketing system: "it issues tickets for read and write operations (128 of each, by default), after which point new read or write operations will queue." The book gives the exact fields to watch for exhaustion — wiredTiger.concurrentTransactions.read.available and wiredTiger.concurrentTransactions.write.available in serverStatus output — reaching zero means new operations of that type are now queuing. This queue depth is the modern equivalent of the old lock-percentage alarm: "the queue size should be low. A large and ever-present queue is an indication that mongod cannot keep up with its load."
Backup approach 1: logical dump (mongodump/mongorestore)
mongodump reads every document through the query layer and writes it out as BSON, organized into per-database, per-collection files. It's the most flexible option — "a good way to back up individual databases, collections, and even subsets of collections" — but the book is explicit about its costs relative to the other methods: "it is slower (both to get the backup and to restore from it) and it has some issues with replica sets." The core problem is that a plain mongodump is not point-in-time: it reads the server while writes keep happening, so a long-running dump can capture an inconsistent mix of before/after state — "you might end up with a situation where user A begins a backup that causes mongodump to dump database A, but while this is happening user B drops A." The book's fix is the --oplog flag on a replica-set-enabled mongod, which records every op that occurs during the dump so mongorestore --oplogReplay can replay them and land on a genuinely consistent snapshot.
Two hard constraints are worth calling out because they're easy to violate by accident. First: never combine fsyncLock() with mongodump — the book's own warning is direct: "mongodump may hang forever if the database is locked." Second, unique indexes complicate a dump/restore cycle because "unique indexes require that the data does not change in ways that would violate the unique index constraint during the copy" — a plain (non---oplog) dump/restore of a collection with a unique index (other than _id) can restore data that momentarily violates its own constraint mid-load.
Backup approach 2: filesystem/volume snapshot
The other family of technique backs up the disk, not the documents: LVM (or an equivalent cloud block-storage snapshot) captures a point-in-time, copy-on-write image of the volume MongoDB's data files live on. The book's summary of the trade: "These methods complete quickly and work reliably, but require additional system configuration outside of MongoDB." Consistency is the catch — "the database must be locked and all writes to the database must be suspended during the backup process" — achieved with db.fsyncLock() before the snapshot and db.fsyncUnlock() immediately after, since "mongod's first action" on a fresh initial sync "is to delete it all" if pointed at the wrong data directory, and a snapshot taken mid-write is not guaranteed consistent. Because WiredTiger's data files "reflect a consistent state as of the last checkpoint" (checkpoints occur every minute), a snapshot doesn't need application-level coordination beyond the fsync lock — the storage engine's own checkpointing does the rest.
The operational shape of this method — lvcreate --snapshot, mount the snapshot, copy or dd | gzip it elsewhere — is fast enough that the book recommends running it against a secondary, not the primary, for exactly the same reason it recommends mongodump be run off-primary: "making a backup can cause strain on a system: it generally requires reading all your data into memory," and that load belongs anywhere except the node serving live writes.
Choosing between them, and the cluster-level picture
For a single server or a replica set, filesystem/volume snapshots and data-directory copies are the book's preferred default — "a filesystem snapshot or data file copy is recommended" — with mongodump positioned as the fallback for finer-grained (single database/collection) backups. For a sharded cluster, there is no way to get a perfectly consistent whole-cluster snapshot while the cluster is live — "sharded clusters are impossible to 'perfectly' back up while active" — so the practical approach is backing up the pieces (each shard's replica set, plus the config servers) with the balancer stopped, or leaning on a managed backup service (Cloud Manager, Ops Manager, or Atlas) that automates exactly this multi-piece coordination.
Book vs. today
Atlas, not Ops Manager, is now the default recommendation for most users. The book already noted Atlas's monitoring "are very similar" to Ops Manager's graphs, but current MongoDB documentation goes further: Atlas is positioned as the primary, actively-developed monitoring and backup surface — with a Performance Advisor, Real-Time Performance Panel, Query Profiler, and built-in alerting — while Ops Manager remains the supported path specifically for self-managed, on-premises deployments that cannot move to a managed cloud service. The self-hosted-first framing of this chapter (Ops Manager as the primary example, Atlas as "similar") has effectively inverted.
Cloud Manager, the book's other named tool, is being wound down for older MongoDB versions. MongoDB's current documentation states Cloud Manager stopped supporting automation, backup, and monitoring for MongoDB 3.6 and 4.0 after August 30, 2024, and directs users toward Atlas (for cloud deployments) or Ops Manager (for on-premises). Cloud Manager was not a book-era concept described as legacy — it's now explicitly framed as the tool to migrate away from.
Atlas's own backup terminology has consolidated around "Cloud Backups." What the book calls "continuous backups" and "cloud provider snapshots" as two Atlas Cloud options are now unified under Atlas's Cloud Backup feature, which snapshots using the native functionality of the underlying cloud provider (AWS, Azure, or GCP) with built-in cross-AZ redundancy, and can additionally replicate snapshots and oplogs across regions for disaster recovery.
mongodump --oplogon a sharded cluster is a harder no than the book implies. The book already warns that from MongoDB 4.2 onward, "you cannot use either mongodump or mongorestore as a strategy for backing up a sharded cluster" for atomicity reasons. Current documentation sharpens this specifically for--oplog: it "has no effect when running mongodump on a mongos instance" — the flag doesn't degrade gracefully, it silently does nothing, so a script assuming point-in-time consistency for a sharded dump is wrong even before the 4.2 atomicity caveat applies. Balancer suspension, cross-shard transaction pauses, and DDL pauses are now explicit prerequisites documented for anymongodumprun against a sharded cluster.
Trade-offs
mongostat/mongotop/serverStatuscost nothing to reach for but require a human watching; Atlas/Ops Manager cost setup time but alert on their own. The raw commands are always available and need no agent installed, which makes them ideal for an active incident. They don't page anyone at 3 a.m. — that requires the dashboard-and-alerting layer the book spends most of the chapter demonstrating via Ops Manager graphs.- Resident memory and page faults are cheap signals with no universal threshold. They tell you a lot about a specific workload once baselined, but the book is explicit that a given fault rate is fine or catastrophic depending entirely on your disk's capacity and your application's latency tolerance — the number is only actionable relative to your own history, not as an absolute.
mongodump/mongorestoretrade backup granularity for speed and correctness risk. They're the only method that lets you selectively back up (or restore into) a single database or collection, and they work with zero system-level setup — but they're slower than a filesystem snapshot, need the--oplogflag for point-in-time consistency on a replica set, and are outright unsupported for whole-sharded-cluster consistency.- Filesystem/volume snapshots trade backup flexibility for speed and simplicity. They're fast, storage-engine-consistent by construction (thanks to WiredTiger checkpoints plus
fsyncLock), and require no query-layer read of the data — but they need system-level tooling (LVM or a cloud block-storage snapshot feature) outside MongoDB's control, and they back up everything on the volume, not a selected subset. - A larger oplog buys resync avoidance at a fixed, low disk cost. The book's two-to-three-day sizing guideline is a straightforward trade: oplog storage "take[s] up basically no memory," so erring toward a longer oplog is nearly free insurance against a lagging secondary going stale during planned maintenance or an unplanned outage.
- Managed backup services (Atlas Cloud Backup, Cloud Manager, Ops Manager) trade operational control for coordinated, cluster-wide consistency. Self-run
mongodump/snapshot scripts give full control over exactly what's captured and when, but coordinating that correctly across every shard and the config servers of a sharded cluster is exactly the multi-piece problem a managed backup service exists to automate away.
Documentation Links
- Shannon Bradshaw, Eoin Brazil, and Kristina Chodorow, "MongoDB: The Definitive Guide", 3rd Edition (O'Reilly, 2020) — Chapter 22-23, "Monitoring MongoDB" and "Making Backups", p. 425-447
- MongoDB Documentation — Monitor for and Improve Performance
- MongoDB Documentation — mongostat
- MongoDB Documentation — mongotop
- MongoDB Documentation — dbCommand serverStatus
- MongoDB Documentation — Atlas Cloud Backup Overview
- MongoDB Documentation — mongodump
- MongoDB Documentation — Back Up a Self-Managed Sharded Cluster with Database Dumps
- MongoDB Documentation — Cloud Manager Overview
- MongoDB Documentation — Ops Manager Overview