Six percent. Still behind.
The first search through a 50 GB log takes ripgrep 54.9–55.8 s and takes us 59.0 s. When this series started, the same comparison read 205 s against 56 s.
We closed a 3.7x gap down to 6% — and then stopped. There was no faster way left to find.
Up front: twelve parts, one idea, and the finish line was the drive itself
Across twelve fixes, searching 50 GB with the free uvf (cold) went from 205 s to 51.70 s; opening the file in the GUI from 100.6 s to 50.44 s; opening, searching and listing numbered results from 189.5 s to 53.7 s; searching a gzip without extracting it from 35.65 s to 17.20 s. Effective read throughput went from 486 MB/s to 967 MB/s — which is exactly what this external USB SSD delivers on a raw read (950–970 MB/s). We hit the ceiling and stopped there. Meanwhile the first query with uvp is still 6–7% slower than ripgrep (59.0 s vs. 54.9–55.8 s, v1.6.6.1), and that one is the price of carrying an index — it does not go away without changing the design. Not one part introduced a cleverer algorithm. Every part did the same thing: count how many times we were reading the same file, then read it fewer times. Mac (Apple M4, 10 cores), 32 GB RAM, external USB SSD, OpenStreetMap XML, cold after sudo purge. Details at the end.
About this series. "First Light" is the pet name of UwView v1.6.6 — the release where both the GUI and the command line
started opening and searching a huge file in a single pass. This series followed what that took, one fix at a time.
Where we started: our own copy said "this is not a speed tool"
As Part 1 described, the v1.6.0 description text contained the sentence "this is not a speed tool." That was not modesty. It was accurate.
- Searching 50 GB with
uvf: 205 s. ripgrep: 56 s. 3.7x behind - Opening 50 GB in the GUI: 100.6 s. klogg: 52.55 s. 1.9x behind
- Open, search, and get a numbered result list end to end: 189.5 s. klogg: 108.14 s
- Reading while building the index: 486 MB/s, on a drive that reads 950–970 MB/s raw — half the media was going unused
Those were the numbers on something we were selling. Twelve parts later, here is what came of them.
Every one of the twelve fixes had the same shape
Laying them out surprised me: not one part introduced a new algorithm. Each was either "read it fewer times" or "read it differently."
- Two and a half passes became one (Part 2). One pass to count lines, one to find matches, half a pass to write output — folded into a single pass: 203 s → 50.82 s
-
-istopped decoding to text (Part 3). Case-insensitive matching folds bytes in place; the one genuine exception is the Kelvin sign U+212A - A required literal is extracted from the regex (Part 4). 13.68 s → 7.75 s, erring toward extracting too little — extracting too much silently drops hits
- The CLI hands its results to the GUI (Part 6). Searching on the command line and then opening the file used to read it twice; now once
- mmap gave way to pread (Part 10). On a sequential read, mmap delivered only 57% of what pread did
- gzip is read without a pipe (Part 8). Calling the OS zlib directly and skipping the 8-byte CRC check: 35.65 s → 17.20 s
That the whole speedup is explained by read counts is the clearest finding of the series. Put the other way round: v1.6.0 was not sloppily written — it was slow by design.
The "make the CPU faster" work never showed up on the clock
By contrast, the part where we made the CPU work faster moved nothing (Part 11).
Counting newlines runs at 1.74 GB/s byte by byte, 3.27 GB/s through the standard search primitive, and 17.9 GB/s with SIMD. In memory that is a 10x spread. Read the same 50 GB off the disk and count it, and every variant lands at 968–969 MB/s.
The media supplies 968 MB/s, so the 10x disappears. Index building went from 486 MB/s to 969 MB/s for a different reason entirely: switching to pread, not a better counting loop.
The numbers
50 GB, cold, after sudo purge
| Against | Metric | Start of the series | v1.6.6 "First Light" |
|---|---|---|---|
| ripgrep | One 50 GB search (uvf) |
205 s vs. 56 s (1/3.7) | 51.70 s vs. 54.94 s (1.06 = level) |
| klogg | Opening 50 GB (GUI) | 100.6 s vs. 52.55 s (1/1.91) | 50.44 s vs. 52.55 s (1.04 = level) |
| klogg | Searching 50 GB (GUI) | 88.9 s vs. 55.59 s (1/1.60) | 53.50 s vs. 55.59 s (1.04 = level) |
| klogg | Open + search + list (end to end) | 189.5 s vs. 108.14 s (1/1.75) | 53.7 s vs. 108.14 s (2.01x) |
| `gzip -dc \ | rg` | Searching a 50 GB-class gzip | not possible at all |
| the drive | Effective read throughput | 486 MB/s | 967 / 966 / 969 MB/s (identical at three sizes) |
One note on wording. We are not calling 1.06, 1.04 or 1.28 a win. By the bar set in Part 9 — a single item under 1.5x is not a difference — all of those are level. Going from 3.7x behind to level is a good outcome; rewriting level as a win would cost every other number its credibility.
What we can call a difference: the 2.01x end to end, and every query after the first. With a .uwvz in place (UwView Pro's compression format plus index), the same 50 GB search returns in 6.56 s, about 8.5x klogg. That clears the bar easily.
The throughput row matters most. At 3 GB, 10 GB and 50 GB it reads 967 / 966 / 969 MB/s — the same number. A figure that does not change with file size is no longer a number about our code; it is a number about the drive. We got to confirm the ceiling by hitting it.
What is still behind
Listing only the ties would be a one-sided picture. Here is where we still lose.
-
The first query with
uvpis 6–7% slower: 59.0 s vs. 54.9–55.8 s at 50 GB (v1.6.6.1). By our own bar that is "level," but we are on the slow side, so we print the number -
A hot (second) plain search at 3 GB: ripgrep 0.32 s vs.
uvf0.54 s. 1.69x behind — that one clears the bar and is a real loss. On small files the fixed startup cost is visible (Part 7) -
-head, which prints only the first matches, is 1.5–6x behindrg | head. ripgrep can read just enough and stop; our index bookkeeping does not let us read less -
3 GB on a 2-core Linux VM is only level (1/1.01). Decompressing the
.uwvzis CPU-bound, so with few cores the advantage of holding an index evaporates (Part 12) - The 258 GB end-to-end GUI run is unmeasured. We have an expectation; we do not publish numbers we have not measured
Why the 6% on the first query does not go away
This is the answer to the opening. It is not slow code. It is the invoice for holding an index.
By Part 12 the first query breaks into three pieces. At 50 GB, cold: building the .uwvz while searching takes 58.6 s, the search that follows takes 0.03 s, total 59.0 s. The older shape — build the index, then re-read the index to search — cost 65.7 s. The 6.8 s of re-reading is gone, folded into building (v1.6.6.1).
Convert the remaining 58.6 s to throughput and you get 828 MB/s. A plain uvf read does 962 MB/s. The gap is compression arithmetic plus writing 5.74 GB.
So: the plaintext is not read twice. It is read once. It is slower because we compress and write while reading it. Removing that cost means dropping compression, and dropping compression removes the 6.56 s second query.
The 6% on query one is what buys the 8.5x on query two. That is why we are not treating it as a bug. If you only ever run one query, plain uvf or ripgrep is the faster choice — and that is the correct way to pick.
What generalizes
(1) Slowness often decomposes into counts. The first response to 205 s was not a profiler; it was counting how many times the file got read. Index 100 s, search 82 s, output 20 s — three separate reads. Once that was on paper, the fix was obvious. Count the passes before you suspect the algorithm.
(2) Measure the ceiling first and you will know when to stop. Knowing the drive reads 950–970 MB/s meant that at 967 MB/s we could declare that direction finished. Without a ceiling you spend weeks chasing the last 3%. The ceiling also protects you from yourself: the "10 GB read cold in 2.65 s" in Part 12 was recognizable as a measurement failure only because the ceiling was known.
(3) Improvements outside the bottleneck do not appear. Counting newlines 10x faster changes nothing when the disk supplies 968 MB/s. Measure where the bottleneck is first. Conversely, on faster media — an internal NVMe — the same change does pay. Whether it pays is a property of the environment, not of the code.
(4) Set the bar for wins before you see the results. Not writing "win" next to 1.06x requires a bar decided before measuring. The bar exists to constrain its author.
(5) Push each remaining loss until it can be stated as a design choice. "Still slower, sorry" and "query one is slower because it is writing the index, which is what buys query two" are different sentences. A loss you can explain that far does not need hiding.
Honest notes
- Every figure comes from one setup: Mac (Apple M4, 10 cores), 32 GB RAM, external USB SSD, OpenStreetMap XML, cold after
sudo purge(rows marked hot are second runs). Results vary by environment - We do not compare seconds across operating systems. Windows and Linux numbers appear only as ratios within each OS (Part 12)
- The "start of the series" column mixes versions (v1.6.0–v1.6.5: GUI figures are v1.6.5, CLI figures v1.6.0–v1.6.2). It is not a snapshot of one release
- Anything under 1.5x is not called a difference. Rows marked "level" could invert in a different environment
- Compared against ripgrep 15.2.0 and klogg 24.11.0.1685. We may not have tuned their settings fully
- The 250 GB / 258 GB end-to-end GUI run is unmeasured, and
uvf … -openat 250 GB is still a v1.6.5 figure, not re-measured on v1.6.6 - The 3 GB "end to end" measurement cannot be taken at all — the file finishes opening while you are still typing the search term — so it is absent from the table
- Match counts agree with klogg (11,274 / 11,393 / 94,979), but that was verified for the main items, not every condition
- I develop one side of this comparison. Discount accordingly
The tools
UwView is a viewer for huge text and log files (free: UwView; paid: UwView Pro).
The same thing from the command line
Everything measured across this series, in one place.
# First query (builds the .uwvz index while searching)
uvp huge.log 'ERROR'
# Second query onward (uses the .uwvz — this is where it pays)
uvp huge.log 'ERROR'
# A plain search with no index (same terms as ripgrep)
uvf huge.log 'ERROR'
# Open the results straight in the GUI (no second read)
uvf huge.log 'ERROR' -open
# Search a gzip without extracting it
uvf huge.log.gz 'ERROR'
# Case-insensitive
uvf huge.log 'error' -i
# Lines only (sed -n p equivalent) / check the exit code
uvf huge.log 'ERROR' -grep
uvf huge.log 'ZZZ-not-found'; echo $?
Exit codes match grep: 0 = found, 1 = not found, 2 = error or incomplete result.
Honestly: the first query with uvp is a little slower than ripgrep, because it builds the .uwvz (UwView Pro's compression format plus index) while searching — at 50 GB, 59.0 s vs. 54.9–55.8 s = 6–7% (v1.6.6.1). It pays off from the second query on: the same 50 GB search returns in 6.56 s, about 8.5x klogg. On small files like 3 GB it is about even. The free uvf is on par with ripgrep at any size (measured article; Mac (Apple M4, 10 cores), 32 GB RAM, external USB SSD, OpenStreetMap XML — one setup, results vary).
Links
All twelve earlier parts.
- Part 1 — the day our own copy said "not a speed tool" (205 s, itemized): https://uvp.y42u.net/en/blog/uwview-fl01-v160-confession-en/
- Part 2 — two and a half passes folded into one: https://uvp.y42u.net/en/blog/uwview-fl02-one-pass-rawgrep-en/
- Part 3 —
-iwithout decoding, and the U+212A exception: https://uvp.y42u.net/en/blog/uwview-fl03-ignorecase-u212a-en/ - Part 4 — regex at fixed-string speed: https://uvp.y42u.net/en/blog/uwview-fl04-regex-required-literal-en/
- Part 5 — three different counts from one command: https://uvp.y42u.net/en/blog/uwview-fl05-deterministic-count-en/
- Part 6 — handing CLI results to the GUI: https://uvp.y42u.net/en/blog/uwview-fl06-cli-to-gui-handoff-en/
- Part 7 — the 3 GB file lost to startup cost, not I/O: https://uvp.y42u.net/en/blog/uwview-fl07-3gb-fixed-cost-en/
- Part 8 — searching a gzip without extracting it: https://uvp.y42u.net/en/blog/uwview-fl08-gz-without-extract-en/
- Part 9 — where a difference begins: https://uvp.y42u.net/en/blog/uwview-fl09-difference-threshold-en/
- Part 10 — mmap gave only 60% of the disk: https://uvp.y42u.net/en/blog/uwview-fl10-mmap-vs-pread-en/
- Part 11 — counting newlines faster moved no clock: https://uvp.y42u.net/en/blog/uwview-fl11-newline-vs-media-en/
- Part 12 — one harness on Mac, Windows and Linux: https://uvp.y42u.net/en/blog/uwview-fl12-three-os-harness-en/
- Removing the re-read on query one (v1.6.6.1): https://uvp.y42u.net/en/blog/uvp-v1661-search-while-building-en/
- Bringing the GUI level with the CLI: https://uvp.y42u.net/en/blog/uvp-gui-catches-up-v166-en/
- Source (GitHub): https://github.com/amru195704/UwView
And if huge logs are eating your storage and you want them kept compressed and still searchable, try UwView Pro: a persistent index, search inside the compressed cache, and about 1/9 the storage, with reopening and searching both a step faster (all OSes; one-time purchase or subscription).
From the developer: a list of my apps, Kindle books, and open source is at GitHub: amru195704.












