Getting Hermit to Linux performance (part 1)
This is the first of two articles about one question: why does a unikernel that does a hundred times less than Linux serve HTTP slower than Linux?
This part is about a single core. It ends with the same axum server ahead of Linux on all six loads we measure, on the same host. Part two is about more than one core, where the answer gets less comfortable.
The short version: the size of the kernel was never the limit. The network stack polled instead of reacting to events, and it walked every socket on every turn. The changes are nine commits in hermit-os/kernel and three in smoltcp-rs/smoltcp, and this article goes through them one at a time.
| Load (ab) | Hermit, before | Hermit, part 1 | Linux 6.12 | Hermit / Linux |
|---|---|---|---|---|
| 1 connection, keep-alive | 296 | 5,364 | 4,604 | 1.16x |
| 8 connections, keep-alive | 1,466 | 6,788 | 4,721 | 1.44x |
| 16 connections, keep-alive | 1,420 | 7,118 | 5,890 | 1.21x |
| 64 connections, keep-alive | 999 | 6,834 | 5,902 | 1.16x |
| 16 new connections per request | 438 | 1,530 | 1,093 | 1.40x |
| 64 new connections per request | 355 | 1,415 | 1,206 | 1.17x |
Requests per second, median of three runs, zero failed requests. All three columns were measured in one session, alternating the three guests at every repetition, so they are comparable with each other. The Hermit build is the tip of the branch linked at the end.
The test setup
- Host: MacBook Pro, Apple M5 Max, macOS. QEMU 11.0.3 in TCG mode, with user-mode networking (slirp). Both guests use 1 vCPU and 256 MiB.
- Application: an axum 0.8 hello-world with a
tokiocurrent-thread runtime. One route returns a 41-byte body. The code is the same for both guests. - Hermit guest: Hermit 0.13.2,
virtio-net-pciwithdisable-legacy=on, the hermit-os forks ofsocket2andtokio. - Linux guest: Alpine
linux-virt6.12.81, a static musl binary asinit,virtio_netfrom the same kernel package. - Load generator:
abon the host.ab -kfor keep-alive,abwithout-kfor one new connection per request.
We wrote nothing in the application to make Hermit faster. The application uses the stock TcpListener::bind from tokio, which asks the kernel for a listen backlog of 1024.
What these numbers are, and what they are not
Everything here comes from QEMU in TCG mode, with user-mode networking, on a laptop. That matters in three ways, and we learned how much only later.
One vCPU. All of it is a single core. Giving both guests more cores does not extend this story, it changes it, and that is part two.
The emulator is inside the measurement. TCG makes an atomic instruction and an MMIO write far more expensive than they are on hardware. Several of the fixes below attack costs that emulation inflates. The costs are real, but their weight relative to each other would be different on a machine with KVM and a real NIC.
The host drifts. The same unmodified binary measured 5,170 requests per second in one batch and 8,098 in another, a few hours apart, with no code between them. So every number in this article comes from runs that alternate the builds being compared, back to back, in one batch. Numbers from two different batches are not comparable, and that includes comparing the per-commit results below with the table at the top.
Before anyone uses this to choose a stack: run it again on a Linux host with KVM and a tap device. We have not, and part two shows why that is the honest next step rather than a detail.
The journey
The first comparison was wrong twice
In August we measured Hermit at 3 times the throughput of Linux, with ab -n100 -c1. That number was wrong for two reasons.
First, the Linux guest used an emulated e1000 card, because our KeyOS port used that card. Hermit used virtio-net, which is paravirtual. With virtio-net on both sides, Linux went from 298 to 1,041 requests per second on the same test. The 3x was the network card.
Second, ab -c1 without keep-alive opens one TCP connection per request. No server is used like this. With keep-alive, the number that matters, Hermit served 300 requests per second and Linux served 4,919.
Lesson: make sure that both guests use the same device model, and measure the load that your users produce.
The rule for the day: measure before you change
We started with several ideas about the cause. Most of them were wrong. The ideas that were wrong, and that we measured as wrong, are listed at the end of this article. Please read that list before you try them again.
The method that worked was simple. We put counters and microsecond timestamps in the kernel and in smoltcp, we printed them at shutdown, and we changed only what the numbers showed. Each change was measured with alternating runs of the old and the new build.
What the numbers showed, in order
- A
writevwith two slices made two TCP segments. The second segment waited for the delayed ACK of the peer. This cost 2.5 ms per request. - The interface was polled on every turn of the wait loop, and every poll walked every socket. With a backlog of 128, that is 131 sockets per turn.
- Sockets that could not send anything were scanned like the others.
- Established sockets with keep-alive never left the scan, because their keep-alive timer is a deadline and not "nothing to do".
- The readiness poll of the listener walked all 128 backlog sockets, twice per request. This was 82 µs of a 278 µs request.
Each of these is one commit below.
Changes to smoltcp, commit by commit
Repository: our fork of smoltcp-rs/smoltcp, branch hermit-perf, based on the tag v0.12.0. The 595 library tests pass after each commit. The changes only affect the iface module.
smoltcp 4d84023: skip sockets a scan cannot make progress on
Problem: socket_egress asks every socket what it wants to send, on every poll. process_tcp walks every socket for every segment that arrives, to find its owner. Both loops read the state of each socket, so each visit pulls a different socket through the cache. A deep listen backlog makes these loops long, and almost all of the sockets in it are idle.
Change: two hints in Meta. They sit next to the handle, not inside the socket. A scan reads them and passes over the socket without a read of the socket:
egress_pending: cleared when a socket reportsPollAt::Ingress. The egress scan andInterface::poll_atskip it.tcp_listen_only: set on a TCP socket seen in theListenstate. Such a socket rejects every segment that carries an ACK, soprocess_tcpskips it for those segments.
Everything that can give a socket work sets the hints again, SocketSet::get_mut first of all. A stale hint costs a needless look. It never causes a missed delivery. Only TCP reports egress_pending. The other socket types are scanned as before.
Result: keep-alive, 1 connection, backlog 1024: 2,846 to 3,354 requests per second. 8 connections: 3,266 to 4,542.
smoltcp 4d66aa7: keep the sockets worth an egress scan on a list
Problem: with the hints, the scans still walk the whole set to read the hint of each socket. For a backlog of 128, that is 131 strides through memory to find the 3 or 4 sockets that can act.
Change: the sockets whose egress_pending holds are on a singly linked list through the set (Meta::egress_next, SocketSet::egress_head). socket_egress and Interface::poll_at follow this list. SocketSet::add puts a socket on the list, remove takes it off, and get_mut puts it back.
This commit also adds SocketSet::peek_mut. It returns a mutable socket without the mark that get_mut sets. Use it only to read state or to register a waker. Hermit needed it. Its readiness poll walked the backlog through get_mut. So every backlog socket went back on the list on every turn of the runtime.
Result: sockets visited per poll went from 131 to 4. Keep-alive, 1 connection: 3,296 to 4,016. 8 connections: 4,439 to 4,961.
smoltcp 9243192: remember when each listed socket next has something to do
Problem: an established TCP socket with keep-alive never reports PollAt::Ingress. Its keep-alive timer makes it report PollAt::Time(t) with t far away. So it never leaves the list, and it gets a full dispatch on every poll. With 64 connections, that was 23 full dispatches per poll, for sockets that all waited on a timer.
Change: after each dispatch, the result of the socket's own poll_at is stored in Meta::egress_at. The egress scan and poll_at skip a listed socket whose time has not come, and they read only the meta to decide. The value is cleared everywhere egress_pending is set. A pending neighbor lookup still overrides it, as before.
Result: keep-alive, 64 connections, alternating runs: 3,921, 3,985 and 4,134 before; 5,583, 5,772 and 5,824 after.
Changes to the Hermit kernel, commit by commit
Repository: our fork of hermit-os/kernel, branch hermit-perf, based on the tag v0.13.2. The kernel builds after each commit, with the feature set of hermit-rs 0.13.2. Commits 5 to 9 need the smoltcp branch above.
hermit-os/kernel 4b0da94: executor, scheduler: fire timers while a task waits with a timeout
This is the patch offered in hermit-os/kernel#2639, from August. Two independent causes made a timed wait last until unrelated I/O woke the core:
block_ongaveTaskNotify::waita wall-clock deadline where it expects a relative timeout.BlockedTaskQueue::addarmed the one-shot timer with the deadline of the added task, even when earlier deadlines sat before it in the list.
The other commits stand on this one because the benchmark app uses tokio::time.
hermit-os/kernel 90acf79: syscalls, fd: write all of a writev in one call
Problem: sys_writev called fd::write once per iovec. Each of those calls is its own block_on and its own interface poll. hyper hands the head and the body of a response to write_vectored as two slices, because tokio reports vectored writes. So each response left as two segments, and the second one waited 2.5 ms for the delayed ACK of the peer.
Change: a new ObjectInterface::write_vectored. The default writes the slices one after the other. The TCP override fills one smoltcp send from every slice. sys_writev makes one call. The old loop returned the error of a later slice after earlier slices had gone out. The default now reports what was written.
Result: keep-alive, 1 connection: 271 to 1,776 requests per second. This is the largest single gain of the day.
hermit-os/kernel 5a1b653: fd: make TCP_NODELAY switch Nagle off
Problem: setsockopt gave the option value straight to set_nagle_enabled, and getsockopt returned nagle_enabled as the option value. Both were inverted. A request for TCP_NODELAY turned Nagle on. The same inversion is on main today.
Change: negate the value in both directions. We confirmed the inversion at the bench before the change: set_nodelay(false) was what removed a Nagle stall.
hermit-os/kernel 7bc2269: executor: poll the interface only when it can make progress
Problem: network_run wakes itself on every turn, and block_on runs it on every turn of its wait loop. A waiting thread polled the whole interface thousands of times per second, and each poll walked every socket. With the 1024 sockets that listen(1024) creates, that was 1024 state machines per turn.
Change: NetworkInterface::poll_if_due polls only when the device holds a frame or when smoltcp's own deadline has passed. Both tests are O(1). A socket that is handed out clears the deadline, because the caller may queue data.
Result: keep-alive, 1 connection, backlog 128: 1,855 to 2,704.
hermit-os/kernel b79c459: net: build against the smoltcp branch with scan hints
The kernel Cargo.toml points smoltcp to branch hermit-perf of our smoltcp fork, commits 4d84023, 4d66aa7 and 9243192. The next commit needs SocketSet::peek_mut.
hermit-os/kernel c4fe154: fd: read socket state and park wakers without marking the socket
Problem: the readiness poll and accept walked the backlog through get_mut_socket. get_mut must assume that the caller queues data, so every backlog socket went back on the egress list on every turn of the runtime. The next poll dispatched them all again: 45 dispatches per poll for one busy connection.
Change: peek_mut_socket, which uses SocketSet::peek_mut, in the three places that only read state or register a waker.
Result: dispatches per poll fell from 45 to 3. Keep-alive, 1 connection: 3,296 to 3,642.
hermit-os/kernel 60f0555: fd: answer listener readiness from a flag instead of a backlog walk
Problem: every sys_poll walked the whole backlog of the listener: a lock, a state read and two waker registrations per socket. That is 29 µs for a backlog of 128. tokio polls twice per request. Timestamps along the request path showed 82 µs of a 278 µs request in this walk.
Change: a ListenNotify waker per listener, registered on every backlog socket in listen() and each time accept() refills the pool. When a backlog socket changes state, smoltcp calls this waker. It raises a flag and wakes the task. The readiness poll answers Pending without a walk while the flag is down. It keeps the flag up while the backlog answers ready, so that a waiting connection is not lost. A non-blocking accept answers EAGAIN without a walk when the flag is down.
Result: keep-alive, 1 connection: 4,105 to 5,415. New connections at 16 clients still all get accepted.
hermit-os/kernel 9831f66: executor: poll the device while a call waits instead of backing off
Problem: a waiting call spun through the exponential backoff of crossbeam. Under emulation, that is about 2,000 pause exits. The NIC interrupt stayed on, so each frame that arrived cost an interrupt on top of the poll that found it. And the loop polled the future on every turn, which is a walk over every fd it waits on.
Change: the NIC interrupt stays off while the loop polls the device itself. The loop spins for a bounded time (250 µs) before it parks. It polls the future only when its waker has fired (TaskNotify::take_wake). The interrupt comes back on before the task parks. Then the loop looks at the receive queue once more. A frame that landed before that point raised no interrupt.
Result: on top of the previous commit, alternating builds: 1 connection 5,415 to 5,697; 8 connections 6,554 to 7,094.
hermit-os/kernel 99fc1c4: apic: do not reload a one-shot deadline the timer already holds
Problem: every return from block_on arms the network deadline again. That is two MSR writes, which the emulator turns into device accesses. Between two syscalls of one request the deadline has not moved.
Change: remember the loaded value, tagged with the core id. Skip a reload of the same value. The timer interrupt clears the memory, so a deadline that has fired is loaded again.
Result: measured together with the previous commit.
Where the time goes now
The timestamps for one keep-alive request at one connection, after all changes, show the fixed costs that remain:
| Item | Per request | Note |
|---|---|---|
recv / read syscalls |
3 × ~9 µs | hyper reads until EAGAIN |
writev syscall |
~28 µs | 4 µs of it is the virtio kick |
write on the tokio eventfd |
~10 µs | the fixed cost of the fd + block_on path |
malloc / free |
14 × 0.5 µs | user-space allocation is a kernel call, and it is cheap |
clock_gettime |
7 × 0.2 µs | |
| user code (mio, tokio, hyper, axum) | ~55 µs |
This breakdown was taken in a different batch than the table at the top, so read the proportions and not the absolute values.\n\nThe fixed cost of one syscall through fd and block_on is about 9 µs, and a request pays it seven times. That is the next target, and part two shows what happened when we went after it.
Ideas that were wrong
We measured each of these. None of them changed the result. Do not try them first.
- A counter in smoltcp so that
acceptskips its scan.acceptis not polled often enough to matter. Removed. - Smaller socket buffers. 1024 sockets × 128 KiB is 128 MiB, but the footprint is not the cost. 4 KiB buffers gave the same numbers.
- Zero spins before the task parks. Slightly worse. The next request arrives during the spin, and a park costs more than the spin.
- A smaller listen backlog. Faster, but wrong: with a backlog of 16, 64 concurrent clients get connection resets. The pool must be as deep as the burst.
- A faster wake of the task from the NIC interrupt. The task almost never parks under load. Its reaction time is not on the critical path.
How to reproduce
The bench harness is in the hermit-axum sandbox. It boots the guest under QEMU, runs ab, and reads the idle CPU of the QEMU process.
Hermit cd hermit-axum
cargo build --release --target x86_64-unknown-hermit
qemu-system-x86_64 \
-cpu qemu64,apic,fsgsbase,fxsr,rdrand,rdtscp,xsave,xsaveopt \
-smp 1 -m 256M \
-device isa-debug-exit,iobase=0xf4,iosize=0x04 \
-netdev user,id=net0,hostfwd=tcp::5577-:80 \
-device virtio-net-pci,disable-legacy=on,netdev=net0 \
-display none -serial stdio \
-kernel boot/hermit-loader-x86_64 \
-initrd target/x86_64-unknown-hermit/release/hermit-axum
Load ab -k -n 1000 -c 1 http://127.0.0.1:5577/ ab -k -n 2000 -c 64 http://127.0.0.1:5577/ ab -n 1000 -c 16 http://127.0.0.1:5577/
Make sure that the Linux guest uses virtio-net-pci with the same flags. Run each point at least three times, and alternate the builds that you compare. The host is noisy: we saw single runs at one third of the normal rate.
What part two covers
One core was the easy half. The second article takes the network off the application thread and onto a core of its own, and most of it is failures worth reading:
- Three bugs that only a second core exposes. Two are in Hermit, one was ours, and each was found by counting rather than by reasoning.
- A park timeout that cost nothing under load and a whole core at idle, which no throughput benchmark would have caught.
- Three well-founded attacks on lock contention between the two cores. All three were measured. All three were slower.
- The measurement that reframes this whole comparison: above two vCPUs both guests are bound to about two cores, and the order between them is decided by the emulator as much as by the two stacks.
Links
- smoltcp-rs/smoltcp, our fork, branch
hermit-perf: commits 4d84023, 4d66aa7 and 9243192 onv0.12.0. - hermit-os/kernel, our fork, branch
hermit-perf: commits 4b0da94 to 99fc1c4 onv0.13.2. - Timer issue with the reproduction crate: hermit-os/kernel#2639.