btop's broken lock
Two threads could hold the same lock.
btop had been crashing every now and then on my machine. In January 2025 I opened issue #1012 with a core dump. The crash seemed related to CPU cores being off-lined, but I could rarely reproduce it. While investigating, I found other sanitizer failures and reported them in #1042. I left it there for a while.
About a year later, I picked it back up. I was tracing shared accesses between the UI and runner threads when I got to btop's custom atomic lock. Turns out, two threads could hold it at once. The waits on Runner::active were broken too. Callers relied on them to coordinate access to shared state, but the waits used relaxed ordering.
Config::current_preset led into Runner::active and the custom atomic helpers.a7d27a6 landed the locking fixes in btop.The lock let a second thread in
atomic_lock wrapped a std::atomic<bool> in RAII: the constructor set the flag to true, and the destructor cleared it. With wait=true, construction was supposed to wait until compare-exchange could change the flag from unlocked to locked.
atomic_lock::atomic_lock(atomic<bool>& atom, bool wait) : atom(atom) {
if (wait)
while (not this->atom.compare_exchange_strong(this->not_true, true));
else
this->atom.store(true);
this->atom.notify_all();
}
atomic_lock::~atomic_lock() noexcept {
this->atom.store(false);
this->atom.notify_all();
}
The loop keeps passing not_true, initially false, as expected:
atom.compare_exchange_strong(expected, desired)
If atom == expected, the operation stores desired and returns true. Otherwise, it returns false and writes the value it observed into expected. The loop reused that changed value on its next attempt.
bool expected = false;while (!atom.compare_exchange_strong(expected, true));Thread A changes the flag from false → true and enters. B starts with expected == false, so its first CAS fails against the flag A has set. That failure changes B's expected to true, and the loop leaves it there.
On B's next attempt, both the flag and expected are true. The CAS succeeds with a true → true transition, and B enters while A still holds the lock. Anything using atomic_lock(..., true) could end up with two threads inside the supposedly protected code.
Where btop was using it
In May 2026, at least two callers used the waiting form:
term_resize() used atomic_lock lck(resizing, true) to prevent concurrent resize handling. The broken CAS meant re-entry was still possible.
Config::write() waited on writelock and then used atomic_lock lck(writelock, true). Two writers could pass the guard at once.
The waits were broken too
btop's runner thread collects data and draws the UI. The main/input thread also changes config and UI state, so the two threads coordinate through Runner::active. The runner marks itself active while using shared state, and main-thread code waits for it to finish before accessing that state.
void atomic_wait(const atomic<bool>& atom, bool old) noexcept {
atom.wait(old, std::memory_order_relaxed);
}
void atomic_wait_for(const atomic<bool>& atom, bool old, uint64_t wait_ms) noexcept {
while (atom.load(std::memory_order_relaxed) == old && ...)
sleep_ms(1);
}
Callers saw Runner::active == false and went ahead with state the runner had just been updating. Relaxed ordering kept the flag coherent, but didn't synchronize those ordinary reads and writes.
My first fix added release ordering to the unlock stores and acquire ordering to the waits. When the acquire observes the release, the runner's earlier writes become visible to the waiting thread.
88b0ed6, then tightened by 5aaca91 and 312592fview commitvoid atomic_wait(const atomic<bool>& atom, bool old) noexcept {
atom.wait(old, std::memory_order_acquire);
}
bool expected = false;
while (!atom.compare_exchange_strong(
expected,
true,
std::memory_order_acquire,
std::memory_order_relaxed
)) {
expected = false;
}
// unlock
atom.store(false, std::memory_order_release);
A failed CAS can use relaxed ordering because only successful acquisition needs to synchronize with the previous unlock. Each retry resets expected to false; success acquires the lock, and unlocking releases it.
How I got here: current_preset
I reached these helpers while tracing accesses to Config::current_preset, a std::optional<int> shared by input handling and drawing.
The runner read it while drawing:
!Config::current_preset.has_value()
? "*"
: to_string(Config::current_preset.value())
Pressing p/P could change the same optional from the input thread. I found an atomic_wait(Runner::active) in that path, but it ran after the reads and mutations it needed to protect:
557fbe5btop_input.cppconst auto old_preset = Config::current_preset;
// reads and writes current_preset here
...
atomic_wait(Runner::active);
Config::apply_preset(...);
current_preset raceMain / input thread
Runner thread
In 557fbe5, I moved the wait before the first access. Two menu paths could also reset current_preset while the runner was drawing; 677336f added guards to those paths.
Timed waits for Runner::active
Some callers also needed a timeout, which C++ atomic wait/notify doesn't provide. The old atomic_wait_for polled the flag, sleeping 1 ms between checks.
I replaced that polling with atomic_waiting_lock, backed by a std::mutex and std::condition_variable. The condition variable handles the timeout directly:
class atomic_waiting_lock {
bool value{};
mutable std::mutex mtx;
mutable std::condition_variable cv;
...
};
void atomic_waiting_lock::wait_for(bool old, uint64_t ms) const noexcept {
std::unique_lock lock{mtx};
cv.wait_for(lock, std::chrono::milliseconds(ms),
[this, old] { return value != old; });
}
Runner::active moved from atomic<bool> to this class, and the runner now obtains an RAII guard through active.lock(). The class handled the waits, timeouts and notifications; the guard handled the lock's lifetime.
I kept the generic atomic_lock, with its corrected CAS loop, for the simpler boolean guards elsewhere in btop.
Other fixes in the PR
Teardown from the signal handler
The SIGINT handler could call clean_quit(0) directly. That function joins threads, writes config, logs, formats strings, accesses the terminal and can allocate memory. Several of those operations are unsafe inside an asynchronous POSIX signal handler.
In 3843042, I changed the handler to set state and call Input::interrupt(). That helper calls kill(getpid(), SIGUSR1), waking normal program flow so it can perform the shutdown outside the handler.
An Intel GPU failure-path leak
The same sanitizer run found leaks in Intel GPU initialization. 327f695 added cleanup for gpu_device_name on two early-return paths and for the engine structure when PMU initialization failed.
The sequence of fixes
557fbe5current_preset access.677336fcurrent_preset.88b0ed63843042clean_quit from SIGINT.327f6955aaca91312592fexpected local, reset it on failure, and notify one waiter.2c631e4The lock's history
In October 2021, 804fe60 replaced atomic-bool spinlocks with mutexes to address a rare deadlock. Three days later, 1601422 switched back to custom atomic-bool locks. The compare-exchange implementation I found in 2026 dates from around that switch; wait/notify came later.
Why it could keep working anyway
- The compare-exchange bug required contention on a
wait=truepath. A thread arriving after the previous owner had unlocked could acquire normally. - The ordering bug affected ordinary shared accesses that relied on the atomic transition for a hand-off.
- CPU count, terminal activity, I/O, compiler optimization, sanitizers and kernel scheduling all changed the timing.
- x86's relatively strong memory ordering could hide symptoms, while the C++ data race remained undefined behavior.
- Even when two threads entered together, a visible failure depended on which state they accessed during the overlap.
TSan reported conflicting accesses without a valid happens-before relation, including runs where btop appeared to work. Those reports gave me the accesses to trace back through the lock and the waits.
The original crash is still open
PR #1649 merged on 23 May 2026 and fixed the #1042 races. The intermittent CPU-hotplug crash from #1012 is still open. If you can reproduce that one, a core dump or a set of steps would help.