Columnar Storage in Go: Fast Financial Aggregation

Consider a financial analytics platform serving real-time portfolio metrics to thousands of clients. Traditional row-oriented storage loads entire transaction records into memory—customer ID, timestamp, instrument, quantity, price, fees—even when clients only request daily trade volumes. At 10 million transactions per day with 50-byte records, this means loading 500MB to calculate a single sum. By switching to columnar storage where each field lives in its own contiguous array, the same aggregation touches only 40MB (the price column), fits in CPU cache, and completes 18× faster. The architecture shift isn’t just about memory efficiency. It’s about aligning data layout with how modern CPUs actually process numerical operations. ...

March 26, 2025 · 21 min · 4400 words · Svein Erik

Neural Networks from Scratch in Rust

In 2017, a fraud detection startup discovered their Python-based neural network inference was creating a hidden cost: 200 milliseconds of latency per transaction. At their scale—15,000 transactions per second—this meant holding $3 million in pending transactions at any moment, exposing them to market risk and regulatory scrutiny. When they rewrote their inference engine in Rust, latency dropped to 8 milliseconds—a 25× improvement—and throughput increased enough to handle 10× growth without additional hardware. The difference wasn’t algorithmic sophistication. It was understanding how neural networks actually execute on real hardware and choosing a language that exposed rather than obscured those realities. ...

March 18, 2025 · 31 min · 6572 words · Svein Erik

Vector Optimization: SIMD, Cache Lines & Memory Bandwidth

In 2019, a quantitative trading firm discovered their portfolio risk calculation was bottlenecked not by algorithmic complexity, but by memory access patterns. By restructuring their data layout and applying SIMD vectorization, they reduced computation time from 47 seconds to 890 milliseconds—a 53× speedup—without changing a single line of business logic. The difference between naive and optimized vector operations isn’t just academic; it’s the difference between real-time decision making and stale analysis in production systems. ...

March 11, 2022 · 33 min · 6981 words · Svein Erik