The counting did get faster.
We replaced a byte-at-a-time == '\n' loop with one that skips ahead. Measured in memory: 1.74 GB/s became 3.27 GB/s. That is 1.88x.
And the time to open a file did not move. So was the change worthless?
Up front: a 1.88x win in RAM that is worth exactly 0% on disk
Switching newline scanning from a byte loop to IndexOf took the in-memory rate from 1.74 GB/s to 3.27 GB/s (1.88x). Against a real 50 GB file on an external USB SSD, both versions land on 968–969 MB/s — the same number, with the same count of 892,239,125 lines. This change on its own did not move the clock. Index building went from 486 MB/s to 969 MB/s in the same release, but the cause of that was Part 10's mmap → pread, not this. We shipped it anyway, because it is preparation that only pays off on faster media. Mac (Apple M4, 10 cores), 32 GB RAM, external USB SSD, OpenStreetMap XML. Details at the end.
About this series. "First Light" is the pet name of UwView v1.6.6 — the release where both the GUI and the command line
started opening and searching a huge file in a single pass. This series follows what that took, one fix at a time.
The problem: counting newlines is the dullest work you cannot skip
To show a huge text file on screen you need markers: which byte offset does line N start at. Keeping nine hundred million of those is too much memory, so you keep every Nth one.
Building those markers means reading the file end to end and counting \n. Nothing about the content matters. You are looking for one byte value.
v1.6.0 did it the obvious way:
for (int i = 0; i < span.Length; i++)
if (span[i] == (byte)'\n') { /* record the offset */ }
We had suspected this loop for a while. Fifty gigabytes is fifty-one billion iterations. Even a few instructions each ought to add up.
Measure the ceiling before you blame anything
As in Part 10, a suspicion like that means nothing until you measure the ceiling first. So we measured three counting strategies against a buffer in memory, with no disk involved at all:
| Strategy | Rate in memory |
|---|---|
Byte at a time, == '\n'
|
1.74 GB/s |
Span.IndexOf, skipping to the next \n
|
3.27 GB/s |
| Vector (SIMD) comparison | 17.9 GB/s |
The gaps are real: 1.88x for IndexOf, about 10x for vectors.
One prior estimate failed here. An internal note said a byte loop "tops out around 1 GB/s". On this machine it does 1.74 GB/s — noticeably better than we had assumed. We had written off the naive loop without measuring it.
Still, 1.88x is 1.88x. The question is where those 1.88x show up in the clock.
The fix: skip instead of compare
We switched SparseLineIndex to Span.IndexOf (commit 64f80c7) and widened reads from 1 MB to 4 MB.
IndexOf wins because it stops comparing one byte at a time: it compares 16 or 32 bytes at once and extracts the position of the hit. Since a line contains exactly one \n, the dozens of bytes in between can be skipped rather than examined. The usual shape of optimisation — do less work, not faster work.
The edit itself is a few lines. Everything interesting happened after it.
The numbers: from disk, both run at 968 MB/s
We ran the same strategies against a 50 GB file on the external USB SSD:
| Strategy | Against the 50 GB file | Lines counted |
|---|---|---|
| Byte at a time | 968 MB/s | 892,239,125 |
IndexOf |
969 MB/s | 892,239,125 |
The 1.88x is gone. 968 against 969 is not even noise.
The reason is arithmetic. pread in 4 MB blocks delivers 966–968 MB/s (Part 10), so bytes arrive at roughly 0.97 GB/s. One strategy can consume them at 1.74 GB/s and the other at 3.27 GB/s. Both are faster than the supply. Surplus capacity turns into waiting, not into seconds saved.
The identical line counts matter for a different reason: a rewrite meant to change speed did not change the answer. Nine hundred million lines, not one of them off.
So what actually took index building from 486 to 969 MB/s?
Opening 50 GB went from 100.6 s (486 MB/s) to 50.44 s (969 MB/s). It is tempting to bank that 1.99x as the payoff for the SIMD work.
We cannot. The measurement above is exactly why. If a byte loop reaches 968 MB/s when reading from disk, then the counting strategy was never what held 486 MB/s down. What held it down was reading through mmap (ceiling 551 MB/s) plus the gap below that ceiling. The cause was Part 10's switch to pread.
Getting this attribution wrong is not just a matter of honesty in an article. Get it wrong, and the next time the same symptom appears you will go polish the counting loop again. Correct attribution is how you decide what to fix next.
Then why ship it?
We shipped a change that moves no seconds. Two reasons.
(1) Media get faster. An external USB SSD delivers 0.97 GB/s; an internal NVMe is in the multi-GB/s range. The moment supply passes about 2 GB/s, the byte loop's 1.74 GB/s becomes the bottleneck. IndexOf at 3.27 GB/s still has headroom then.
Put another way: a difference invisible at today's measurement becomes visible when the hardware changes. Part 9 set a rule that nothing under 1.5x counts as a difference — but that rule assumes the conditions are fixed. When the conditions themselves are expected to move, something that looks like no difference today can still be worth having.
(2) Spare CPU is spendable elsewhere. While the read is in flight the CPU is idle, and that idle time is a budget: what else can you count during the pass you are already paying for? Part 6's "build the line index as a by-product" and Part 2's "fold counting, matching and writing into one pass" both spend from that budget. Making one stage cheaper means more of the budget is available to everything else.
To be straight about it, we did not go as far as the 17.9 GB/s vector version. IndexOf already has plenty of headroom and on this medium 3.27 and 17.9 are indistinguishable. That decision gets revisited when we have hardware that can tell them apart.
Being straight about this
- There is no speed improvement in this part. At 50 GB it is 968 versus 969 MB/s. It is the one instalment in this series where the number did not move
- A 10 GB file measured 1,026 MB/s, which is above the 50 GB figure of 969. That is because 10 GB partly stays in the page cache on a 32 GB machine, while 50 GB (1.6x RAM) does not. The two are not comparable and we do not put them side by side as a ranking
- The three in-memory rates (1.74 / 3.27 / 17.9 GB/s) are measured with no disk involved. They describe capacity, not runtime
- The older internal note claiming a byte loop tops out near 1 GB/s did not reproduce on this machine (1.74 GB/s). Using a stale estimate as evidence is the mistake this part recorded
- That
IndexOfwill pay off on faster media is a prediction, not a measurement. We have not re-measured on internal NVMe - All figures come from Mac (Apple M4, 10 cores), 32 GB RAM, external USB SSD, OpenStreetMap XML (51,254,526,392 bytes). We do not compare seconds across operating systems
- I write about my own product. Discount accordingly
What generalises
(1) "Faster" means nothing without saying where you measured. 1.88x in memory; 1.00x from disk. Same change, same machine, same day. Both are true, and the writer gets to pick which one to publish. Publishing both is the only honest option.
(2) Attribute the win correctly. When a 2x improvement and a SIMD rewrite land in the same release, one will get credited for the other. Proximity in time is not causation. Separating them required measuring the changes one at a time.
(3) You may ship an improvement that does nothing today — if you can state what has to change for it to matter. Here we can: "once bytes arrive faster than about 2 GB/s". If you cannot state that condition, the rewrite is a matter of taste. Shipping the condition alongside the change is what makes it checkable later.
(4) Always verify the answer, not just the clock. The scariest outcome of a performance rewrite is that it gets faster and returns something different. The line count is compared on every run. Same answer first, then compare speed.
The tools
UwView is a viewer for huge text and log files (free UwView / paid UwView Pro).
Doing the same thing from the command line
# Total line count (one pass over the file)
/usr/bin/time -p uvf huge.log -lines
# Count only (no hit text kept in memory)
uvf huge.log 'ERROR' -count
# Compare against the raw throughput of the medium
dd if=huge.log of=/dev/null bs=4m count=5000
# Confirm you get the same answer before comparing speed
uvf huge.log 'ERROR' -out mine.txt
rg 'ERROR' huge.log > theirs.txt
Exit codes match grep: 0 = found, 1 = not found, 2 = error or incomplete result.
Honestly: the first query with uvp is a little slower than ripgrep, because it builds the .uwvz (UwView Pro's compression format plus index) first — at 50 GB, 59.0 s vs. 54.9–55.8 s = 6–7% (v1.6.6.1). It pays off from the second query on: the same 50 GB search returns in 6.56 s, about 8.4x ripgrep. On small files like 3 GB it is about even. The free uvf is on par with ripgrep at any size (measured article; Mac (Apple M4, 10 cores), 32 GB RAM, external USB SSD, OpenStreetMap XML — one setup, results vary).
Links
- Part 2 — folding two and a half reads into one: https://uvp.y42u.net/en/blog/uwview-fl02-one-pass-rawgrep-en/
- Part 6 — building the line index as a by-product of the pass: https://uvp.y42u.net/en/blog/uwview-fl06-cli-to-gui-handoff-en/
- Part 7 — putting the change into something it should not help: https://uvp.y42u.net/en/blog/uwview-fl07-3gb-fixed-cost-en/
- Part 9 — where a difference begins: https://uvp.y42u.net/en/blog/uwview-fl09-difference-threshold-en/
- Source (GitHub): https://github.com/amru195704/UwView
And if huge logs are eating your storage and you want them kept compressed and still searchable, try UwView Pro: a persistent index, search inside the compressed cache, and about 1/9 the storage, with reopening and searching both a step faster (all OSes; one-time purchase or subscription).
From the developer: a list of my apps, Kindle books, and open source is at GitHub: amru195704.












