
Why one big folder stalls a client
A user opens one folder and the file manager hangs for half a minute. Throughput tests look healthy, the network is quiet, and the array is nowhere near saturated. The culprit is frequently one of the large directory listings the client is trying to enumerate, a workload that stresses metadata and round trips instead of bandwidth. Understanding what happens during an enumeration explains why it stalls and what folder design can do about it.
Enumeration is a conversation, not a transfer
Listing a directory means asking the server for entries in batches. Each reply is capped by a size the client requests, and on NFS the client also supplies a hint for how many bytes of entries it wants back, so a folder with several hundred thousand entries needs a great many request and response pairs before the client has the whole picture. Latency multiplies across that sequence. A link with 1 ms round-trip time and a few thousand sequential requests adds seconds of pure waiting before any rendering starts, regardless of how many gigabits the link can carry. Large directory listings are a latency problem first and a bandwidth problem a distant second.
What NFS does
NFS clients use READDIR to fetch names and, in NFSv3, READDIRPLUS to fetch names together with attributes. The client passes a cookie marking where the last batch ended, plus size hints for how much data to return. READDIRPLUS saves follow-up GETATTR calls when the user will want sizes and timestamps anyway, but it makes each reply heavier, and it wastes effort if the client only needs names. NFSv4 folds the same ideas into a single READDIR operation with a requested attribute set.
Then there is the habit behind most slow listings. Running ls -l forces a stat on every entry, while ls -f or ls -U skips sorting and attribute lookups. A script that does a find over a huge tree can generate far more metadata traffic than the equivalent data copy.
What SMB does
SMB clients send a QUERY_DIRECTORY request, naming an information class and an output buffer size, and the server returns as many entries as fit. The client repeats the request until the server answers that no more files remain. Windows Explorer then compounds the load, because it often requests icons, thumbnails, or extra properties as it paints each row. A folder of ten thousand images can trigger far more work than the listing itself, which is why the same folder can feel fine from a shell and terrible from a desktop. It is also why large directory listings are worth testing with the tool your users actually use.
Why the server side hurts too
Directory layout inside the filesystem matters, and so does the appliance underneath it. Choosing a purpose-built NAS appliance with enough memory for metadata cache and fast media for the metadata itself makes large directory listings far less painful. Modern filesystems use tree or hash structures so lookups stay fast. Even so, a directory with millions of entries still means a lot of metadata to read, and a cold cache turns that into random disk I/O. Sorting is another cost: some clients ask for sorted output and the server may have to produce an order that the on-disk structure does not give for free.
Caching helps when it is warm. Linux NFS clients hold directory attributes for a configurable period, with acdirmin defaulting to 30 seconds and acdirmax to 60, which means a second listing shortly after the first can be nearly instant. That is also why the first morning access is the one people complain about.
Folder design that stays fast
The cheapest fix is structural. Instead of one flat directory holding every file, spread entries across a shallow hierarchy keyed by something stable, such as a date, a customer identifier, or the first characters of a hash. A few hundred entries per directory is comfortable; hundreds of thousands starts to feel slow whatever the hardware. Applications that create files at high volume should be configured to shard from the start, because reshaping a populated tree later means a long migration.
Retention matters as well, because large directory listings tend to grow quietly: nobody plans for a million files, they just never delete the first hundred thousand. Old, finished data can be moved to an archive tree so that active folders stay small. Backup chains are a good example: the per-job files in a NAS backup repository for Veeam grow steadily, so give each job its own directory and expire restore points on schedule instead of letting everything accumulate in one place.
Diagnosing a stall
Start by timing the listing itself with a plain ls -f in a shell on the client, which isolates directory reads from attribute and rendering work. If that is fast and ls -l is slow, attribute lookups are the bottleneck. On the NFS side, nfsstat -c shows counts of readdir and getattr calls, and a packet capture shows the request pattern directly. Compare the same folder from a second client and from the server console to separate network effects from storage effects.
Pay attention to what else is running. Antivirus scanners, search indexers and sync agents all enumerate directories on a schedule, and each pass competes with interactive users. Schedule them outside peak hours where you can, and exclude the busiest trees from real-time scanning if your policy permits it. Tidy folders make every pass cheaper.
Security and housekeeping
Directory sprawl has a security side too. Folders no one remembers are rarely reviewed for access, and a flat share with thousands of owners-by-accident is hard to audit. Pruning and sharding make the review tractable, which fits with the broader NAS security best practices that most teams already follow.
The next time one folder is slow while everything else is quick, look at the number of entries before looking at the network. Fewer entries per directory, sensible caching and listing tools that ask only for names will usually do more than another faster link.
Add comment
Comments