September 29, 2026 • By Brian Marete
I am pleased (or slightly embarrassed) to announce a new version of zetasketch-cpp (https://github.com/rowkeydb-com/zetasketch-cpp), the C++23 distributed counting and cardinality estimation library designed for bit-exact compatibility with Google’s Cloud Bigtable ZetaSketch implementation.
This compatibility is critical for RowKeyDB and useful if you work in C++ and seamlessly create or merge sketches with those used in Cloud Bigtable or BigQuery. I previously wrote about the benefits of this property here: https://rowkeydb.com/blog/zetasketch-hll-distributed-aggregation/
The headline features are:
September 16, 2026 • By Brian Marete
You want an exhaustive proof that your algorithm is correct, but the mathematics of complexity gives you a sobering blow to the head.
Then you have to get devious about how to gain substantial confidence in the part of the state space not covered by the TLC checker due to memory and time intractability.
I continue to make progress with the validation of the code and TLA+ model of RowKeyDB’s distributed consensus algorithm, and I will write a short article soon about the complementary role of:
August 25, 2026 • By Brian Marete
RowKeyDB will be a drop-in replacement for HBase (via the adapter) and Cloud Bigtable (in any cloud or DC environment) at any conceivable scale and uptime requirement, there will be no excuses.
In general, that will require performant and stable and linearly scalable multi-node capabilities, a beta of which will be ready in about 5 weeks.
But even a single-node version that already can achieve tens of thousands of read/writes per second is pretty audacious, if I may say so myself, and deserves an update on progress towards v1.0.0 GA release.
August 13, 2026 • By Brian Marete
I am pleased to announce the release of v0.1.0-beta.5 of the single-node version of RowKeyDB: https://github.com/rowkeydb-com/rowkeydb-releases/releases/tag/v0.1.0-beta.5
The headline feature is new support for a Google Cloud Bigtable bit-exact (fully compatible) HyperLogLog++ cardinality estimation aggregate data type, completing support for all Bigtable MutateRow RPC messages. RowKeyDB already has full support for the read side (ReadRows) including heavily tested, correct and efficient implementation of Bigtable’s filters in their full range of possible combinations.
August 7, 2026 • By Brian Marete
I thought I should say something to illuminate the importance of the bit-exact output of zetasketch-cpp—which I announced in my previous post—with the HyperLogLog++ output of Google Cloud Bigtable, Google BigQuery, etc.
The second cleverest feature of HyperLogLog++ (which is unlocked when bit-exactness is available) is the fast aggregation of distributed counting. Using the merge operation, you can take a “sketch” counting distinct “things” aggregated by Machine A in Service S, and immediately merge it with a sketch counting the same things on a completely different Machine B in Service T. The result is a sketch with the correctly aggregated cardinality estimate.
August 7, 2026 • By Brian Marete
How do you count, say, billions of distinct IP addresses for every visit to your site without gobbling up more memory than you can afford?
I am happy to present zetasketch-cpp (https://github.com/rowkeydb-com/zetasketch-cpp)—an implementation of HyperLogLog++ (HLL++), the cardinality estimation algorithm that solves exactly that problem for you.
As far as I know, this is the first open-source C++ implementation of Google’s ZetaSketch library. Furthermore, I have carefully tested it to produce bit-identical output to Google’s Java ZetaSketch library. I designed the testing to cover a large variety of the algorithm’s state space, including the various complex corner cases of inputs when merging sketches (see the verification details here: https://github.com/rowkeydb-com/zetasketch-cpp#verification-of-bit-exactness-of-output-compared-to-googles-zetasketch).
July 31, 2026 • By Brian Marete
Building RowKeyDB with formal methods and Deterministic Simulation Testing (DST) continues to be quite the fascinating odyssey for me.
DST is a methodology I am using to test that the code implementing the Raft protocol and its six necessary extensions behaves as expected under a wide variety of physical system faults—the kind that happen almost every minute in production.
The space of faults—network partitions (symmetric and asymmetric), delayed packets, CPU starvation events, disk slowness events, clock skew, and their respective frequencies—that can be simulated and tested is exceptionally large. Consequently, managing this state space is not a trivial problem.
June 28, 2026 • By Brian Marete
The C++ implementation of gRPC, otherwise an excellent library, has led me to petty chicanery and sin.
The Go and Java implementations of gRPC have natural, well-documented interfaces for the cheap (before any allocations, before any deserialization of protobuf payloads) interception of every request.
gRPC-Go has the Tap interface, specifically for this kind of thing.
gRPC-Java has the ServerTransportFilter interface for this kind of thing.
But for some reason that I cannot quite figure out, in the gRPC C++ implementation, what would at first glance pass for the natural interface for cheap circuit breaking—the grpc::experimental::Interceptor interface—is not only experimental and deprecated, but it is also not cheap. By the time a request gets there, the payload has already been deserialized, a lot of allocation has already taken place, and in general, a lot of computation has taken place.
June 28, 2026 • By Brian Marete
Girding my loins for the big push of adding multi-node mode to RowKeyDB, I spent last night adding an MSan (MemorySanitizer) CI workflow.
I had been putting MSan off mainly because I dreaded the headache—libstdc++ is infamously essentially un-instrumentable for the purpose, and so I had to swap that out for libc++, like everyone else does, for the MSan build.
In the end, it was not as painful as I expected, and the process was actually fruitful: Speaking a little roughly, libc++ and libstdc++ represent std::chrono::system_clock differently, and the libc++ introduction caused an overflow in clock arithmetic that was not there under libstdc++. Just the same, under MSan, this overflow was caught by a pre-existing unit test, which was fairly satisfying. This and an actual use of an uninitialized variable in observability test code (a real MSan find) were the only obstacles to a clean MSan build and test run.
June 26, 2026 • By Brian Marete
I am happy to report that RowKeyDB passes the entire integration and conformance test suite published for the Google HBase -> Cloud Bigtable adapter (which can be found here: https://github.com/googleapis/java-bigtable-hbase).
The single node RowKeyDB beta binary passed the test suite without modification for BOTH the persistent and memory backends, on the first try.
You do not have to take my word for it. You can download the binary I published a few days ago and test it yourself using these instructions on the release page: https://github.com/rowkeydb-com/rowkeydb-releases#hbase-compatibility-and-conformance-testing
June 25, 2026 • By Brian Marete
When designing RowKeyDB, we knew that a high-performance, Bigtable-compatible API shouldn’t just be restricted to durable storage. Many modern infrastructure challenges require the rich semantics of wide-column stores, but strictly for volatile, low-latency data.
That is why RowKeyDB ships with a native, memory-only backend. Simply pass --storage_backend=memory at startup, and your instance is ready.
It is common for databases to offer an in-memory mode that acts merely as a development crutch. We took a different path. The RowKeyDB memory backend is subjected to the same rigorous correctness tests as our durable engine. It supports the exact same Bigtable API. If your application works against the memory backend, it will work against the durable backend—and vice versa.