AI Trading System Architecture for Financial Markets

Financial markets generate petabytes of data daily: price ticks, order book updates, news feeds, earnings reports, social media sentiment, and macroeconomic indicators. Traditional quantitative finance relies on human-designed models—moving averages, mean reversion strategies, factor models—that capture known patterns but struggle to adapt to regime changes and novel market dynamics. Modern AI agents combine machine learning with systematic trading infrastructure to process multimodal signals, estimate future price movements, and execute trades at scale. ...

November 22, 2025 · 17 min · 3487 words · Svein Erik

Neural Networks from Scratch in Rust

In 2017, a fraud detection startup discovered their Python-based neural network inference was creating a hidden cost: 200 milliseconds of latency per transaction. At their scale—15,000 transactions per second—this meant holding $3 million in pending transactions at any moment, exposing them to market risk and regulatory scrutiny. When they rewrote their inference engine in Rust, latency dropped to 8 milliseconds—a 25× improvement—and throughput increased enough to handle 10× growth without additional hardware. The difference wasn’t algorithmic sophistication. It was understanding how neural networks actually execute on real hardware and choosing a language that exposed rather than obscured those realities. ...

March 18, 2025 · 31 min · 6572 words · Svein Erik

ML Risk Models for Nordic Power Futures & GoO Portfolios

The Nordic power market is one of the world’s most liquid and sophisticated electricity markets, trading over 500 TWh annually across Norway, Sweden, Finland, and Denmark. Power producers, industrial consumers, and financial players manage portfolios worth billions of euros, exposed to extreme price volatility driven by weather patterns, hydroelectric reservoir levels, wind generation variability, and cross-border transmission constraints. A single winter storm can swing prices from €50/MWh to €500/MWh within hours. A mild autumn can crash prices to near-zero as hydroelectric reservoirs overflow. Managing risk in this environment is not optional—it’s existential. ...

December 10, 2024 · 35 min · 7448 words · Svein Erik

LLMs & Foundation Models in Recommenders (Part 6 of 6)

Part 6 of 6 | ← Part 5: Implementation Emerging AI Capabilities The architectures described in this article represent the current state of the art, but fundamental limitations remain unsolved. The next generation of recommendation systems will be shaped by advances in foundation models, multimodal understanding, and content reasoning. Foundation Models and LLMs Current recommendation systems treat content as opaque embeddings. A video is a 256-dimensional vector; the system doesn’t “understand” that it’s a cooking tutorial showing a dangerous knife technique. Large language models change this. ...

July 10, 2024 · 14 min · 2973 words · Svein Erik

Feed Ranking Architecture & Operations (Part 5 of 6)

Part 5 of 6 | ← Part 4: Ethics & Safety | Part 6: Advanced Topics → Case Study: Feed Ranking Architecture The following synthesizes public disclosures from major platforms (Meta, TikTok, YouTube, Twitter/X) into a representative architecture. Request Flow sequenceDiagram participant Client participant Gateway participant Retrieval participant Ranking participant Reranking participant Collator Client->>Gateway: Feed request Gateway->>Retrieval: Get candidates (user_id, context) Retrieval-->>Gateway: 5000 candidates Gateway->>Ranking: Score candidates Ranking-->>Gateway: 500 scored items Gateway->>Reranking: Apply diversity, policy Reranking-->>Gateway: 100 items Gateway->>Collator: Mix organic + ads + notifications Collator-->>Gateway: 50 items Gateway-->>Client: Feed response (paginated) Component Details Component Implementation Retrieval Two-tower model (user/item embeddings) + graph-based (friends’ posts) + trending Ranking Multi-task DCN with 100+ features; outputs P(click), P(like), P(share), P(hide), E(watch_time) Re-ranking MMR for diversity; policy filters; creator frequency caps Feed Collator Interleaves organic posts, ads (from separate auction), and system notifications Latency Budget Stage Target Latency (P99) Feature fetch 10 ms Retrieval 30 ms Ranking (500 items) 50 ms Re-ranking 10 ms Total < 150 ms Caching, precomputation, and model optimization keep end-to-end latency within budget. ...

July 10, 2024 · 8 min · 1539 words · Svein Erik

Recommender Ethics, Fairness & Governance (Part 4 of 6)

Part 4 of 6 | ← Part 3: Production Systems | Part 5: Implementation → Ethical Considerations Recommendation systems shape public discourse and individual well-being. Responsible design requires attention to: Amplification Harms Misinformation: Engagement-optimized systems may amplify sensational or false content. Polarization: Filter bubbles reinforce existing beliefs; users may not encounter diverse perspectives. Addiction: Infinite scroll and personalized feeds maximize time-on-site, potentially at the cost of user well-being. Mitigation Approaches Approach Description Integrity classifiers Demote or remove content flagged as harmful Diversity injection Ensure feeds include diverse viewpoints Time-spent nudges Notify users after extended sessions Transparency Explain why items were recommended (“Because you liked X”) User controls Allow users to tune recommendations, hide topics, or opt out Fairness Recommendation systems can perpetuate or amplify societal biases. Formal fairness metrics provide mathematical frameworks for measuring and mitigating these harms. ...

July 10, 2024 · 10 min · 2062 words · Svein Erik

Recommendation Systems in Production (Part 3 of 6)

Part 3 of 6 | ← Part 2: Ranking | Part 4: Ethics & Safety → Evaluation and Metrics Recommendation systems require rigorous evaluation across offline, online, and long-term dimensions. Offline Metrics Metric Definition Use Case AUC-ROC Area under ROC curve for engagement prediction Pointwise model quality Log-loss Cross-entropy of predicted probabilities Calibration quality NDCG@k Normalized discounted cumulative gain at rank k Ranking quality Recall@k Fraction of relevant items in top-k Retrieval coverage Hit Rate Whether the engaged item appears in top-k Retrieval success Offline metrics use held-out interaction logs; they are necessary but not sufficient for production decisions. ...

July 10, 2024 · 27 min · 5544 words · Svein Erik

Ranking & Re-ranking Recommendations (Part 2 of 6)

Part 2 of 6 | ← Part 1: Architecture | Part 3: Production Systems → Ranking Models The ranking stage scores candidates with a model trained to predict user engagement. Unlike retrieval, ranking models can afford to examine detailed feature interactions. Problem Formulation Ranking is typically framed as pointwise, pairwise, or listwise learning to rank: Approach Loss Function Pros Cons Pointwise Cross-entropy, MSE on engagement labels Simple; scales to large data Ignores relative ordering Pairwise BPR, hinge loss on (positive, negative) pairs Captures preference structure Expensive pair sampling Listwise LambdaRank, softmax over slate Directly optimizes ranking metrics Complex; requires full slate Most production systems use pointwise classification (predicting P(click), P(like), P(share)) due to simplicity and scalability, with calibration layers to combine predictions into a single score. ...

July 10, 2024 · 22 min · 4529 words · Svein Erik

Recommendation System Architecture & Features (Part 1 of 6)

Part 1 of 6 | Part 2: Ranking & Re-ranking → Introduction Social media platforms generate value by connecting users with content they find engaging. The recommendation system—the algorithmic machinery that selects which posts, videos, or accounts to surface—is the engine of this engagement. A well-designed recommender balances multiple objectives: user satisfaction, content diversity, creator economics, and platform health. This article examines the architecture of production-grade recommendation systems, drawing on published designs from major platforms while maintaining a principled, systems-oriented perspective. ...

July 10, 2024 · 24 min · 4944 words · Svein Erik