After working on several quantitative finance and energy market projects, I recently took time to revisit and solidify my understanding of core machine learning concepts. Here’s what stood out—and how these fundamentals connect to real-world applications I’ve built.

The Foundation: Linear Models

Linear Regression & Gradient Descent: The Workhorse of Prediction

Linear regression remains one of the most elegant and powerful tools in ML. While I’ve used it extensively in my energy trading projects, revisiting the mathematics reminded me why it’s so effective: the beautiful interplay between the cost function and gradient descent.

Gradient descent is like hiking down a mountain in thick fog—you feel around for the steepest direction downward and take small steps until you reach the valley. This simple metaphor belies the sophistication of how models learn optimal parameters.

From my projects: In my energy momentum ML trading strategy, I used linear regression as part of the feature engineering pipeline. While Random Forest was the primary classifier, linear features like price momentum and moving average ratios formed the foundation. Understanding how gradient descent optimizes these relationships was crucial for interpreting model behavior.

Real-world impact: Zillow uses regression models to estimate home values. Netflix uses it to predict how much you’ll enjoy a show. Your fitness tracker? Linear regression helps predict your calorie burn.

Logistic Regression: Classification That Powers Real Systems

Logistic regression transforms the linear prediction into a probability using the sigmoid function—that beautiful S-shaped curve that maps any real number to the range [0, 1]. Despite its name suggesting regression, this is fundamentally a classification algorithm.

The key insight? Using log loss instead of mean squared error. MSE creates multiple local minima when applied to classification problems, causing gradient descent to get stuck. Log loss provides a smooth, convex optimization landscape.

Real-world application: In my Nord Pool market surveillance alert system, I used logistic regression as part of the anomaly scoring pipeline. While the final system combined statistical methods with rule-based detection, logistic regression helped classify whether price spreads between zones indicated normal congestion versus potential market abuse—critical for regulatory compliance under REMIT.

Data Engineering & Preprocessing: Where Most of the Work Happens

Revisiting preprocessing fundamentals reminded me that feature engineering and proper data preparation often matter more than algorithm selection.

Feature Scaling: Critical for Convergence

Standardization (subtracting mean, dividing by standard deviation) ensures features operate on comparable scales. Without it, a feature measuring prices in thousands would dominate a feature measuring percentages—even if both are equally important.

The formula is elegantly simple: (x – mean) / std_dev

Critical lesson: Data Leakage

One of the most important concepts I reinforced: never let information from your test set influence model training. This means:

  • Fit your scaler on training data only
  • Apply those same transformations to test data
  • Never standardize train and test together

In my renewable energy portfolio optimization project, this was crucial. When backtesting my Sharpe-optimal wind/solar allocation (52% wind, 48% solar), I had to ensure no future price information leaked into my GARCH volatility forecasts. The out-of-sample Sharpe ratio (1.92) validated that the model generalized well—confirming no leakage had occurred.

Linear Algebra: The Mathematical Backbone

Revisiting linear algebra fundamentals—matrices, vectors, dot products—reminded me why this mathematics is so central to machine learning efficiency.

The key insight: matrix operations enable vectorization. Instead of looping through thousands of examples, we stack them into matrices and compute predictions in a single operation. This is why modern ML can process millions of data points efficiently.

From my trading work: In my energy momentum ML trading strategy, the Random Forest model processes 147 features across thousands of time steps. The efficiency comes from vectorized operations—NumPy’s matrix multiplication handles the entire feature matrix at once, computing predictions for all examples simultaneously. Understanding the linear algebra underneath helps optimize performance and debug issues when they arise.

Ensemble Methods: When One Model Isn’t Enough

Decision Trees & Random Forests

Decision trees make predictions through a series of yes/no questions, recursively splitting data based on feature values. The algorithm selects splits that maximize information gain—the reduction in Gini impurity at each node.

But single trees overfit badly. Enter random forests: train hundreds of trees on bootstrapped samples of your data, then aggregate their predictions. This bagging (bootstrap aggregating) approach dramatically reduces overfitting.

In production: My energy momentum ML trading strategy uses Random Forest for regime classification across 5 energy commodities (WTI Crude, Brent, Natural Gas, Heating Oil, Gasoline). With 147 engineered features per asset, Random Forests handled the high-dimensional feature space better than any single model.

Key results:

  • Strategy outperformed buy-and-hold by +22.81% on average (2023-2024)
  • Best performance on Heating Oil: 16.86% annual return, Sharpe ratio 0.65
  • Random Forest classified market regimes (bearish/neutral/bullish) with ~71% directional accuracy

The bootstrapping mechanism was critical—it prevented overfitting on specific market conditions and made the model robust to regime shifts.

Neural Networks & Deep Learning

Revisiting neural network fundamentals—forward propagation, backpropagation, activation functions—reinforced how these concepts build on everything else: linear algebra (matrix operations), calculus (gradients), and optimization (gradient descent).

Each neuron performs: h = wx + b, then applies an activation function f(h). Chain thousands of these together, and you get networks capable of modeling extremely complex, non-linear relationships.

Key innovations that make it work:

  • Activation functions (ReLU, sigmoid, tanh) introduce non-linearity
  • Backpropagation efficiently computes gradients through the chain rule
  • Automatic differentiation (PyTorch’s autograd) handles the calculus
  • Stochastic gradient descent with minibatches enables training on massive datasets

Why I appreciate the fundamentals more now: While I haven’t deployed neural networks in production yet (Random Forests and GARCH models were more appropriate for my energy/finance projects), understanding the architecture helps me recognize when deep learning would be the right tool—particularly for problems with complex feature interactions or when you have massive datasets (>100k examples).

Unsupervised Learning: K-Means Clustering

K-means clustering finds natural groupings in unlabeled data. The algorithm:

  1. Initialize k centroids randomly (or smartly with k-means++)
  2. Assign each point to its nearest centroid
  3. Move centroids to the mean of their assigned points
  4. Repeat until convergence

The elbow method helps determine optimal k by plotting inertia (sum of squared distances) vs. number of clusters. The “elbow”—where adding clusters yields diminishing returns—suggests the right k.

Practical application: While my portfolio optimization project used Markowitz mean-variance optimization rather than clustering, I’ve applied similar distance-based thinking to energy market segmentation. In ERCOT trading, clustering could identify distinct price regimes (low-volatility, high-volatility, negative-price periods) to inform strategy selection—similar to how my Random Forest classifies market regimes.

The mathematical parallel: both k-means and portfolio optimization minimize a distance metric (k-means: Euclidean distance to centroids; portfolio optimization: variance from target return).

Key Insights Reinforced

1. Algorithm Selection Matters—But Context Matters More

When to use what:

  • Linear/Logistic Regression: Interpretable, fast, works well with smaller datasets (<10k examples)
  • Random Forests: Handles non-linearity, robust to outliers, excellent for tabular data
  • Neural Networks: Best for very large datasets (>100k), complex patterns, image/text/audio
  • K-Means: Exploratory analysis, customer segmentation, data compression

From my projects: Random Forests dominated my energy trading work (147 features, 5 assets, regime classification). For my Nordic power price volatility forecasting, I found GARCH models (statistical time series) outperformed ML approaches. Knowing when not to use ML is as important as knowing how to build models.

2. The Math Builds on Itself

  • Linear algebra → matrix operations power everything
  • Calculus → gradients drive optimization
  • Probability → cost functions, Bayes’ theorem
  • Statistics → hypothesis testing, confidence intervals

The mathematical foundations carry across all quantitative work—whether ML, statistical modeling, or numerical methods.

3. Preprocessing Can Make or Break Your Model

Critical steps I’ve learned the hard way:

  • Proper train/test splitting (no data leakage!)
  • Feature scaling (especially for gradient-based methods)
  • Handling missing data (imputation vs. deletion)
  • Outlier treatment (robust scalers, winsorization)
  • Feature engineering (domain knowledge > raw features)

In my energy trading strategy, I engineered 147 features from raw price data: momentum indicators, volatility measures, moving average ratios, correlation metrics. These hand-crafted features outperformed raw prices by a wide margin.

4. Validation Is Non-Negotiable

My validation framework (adapted from quantitative finance):

  1. Out-of-sample testing: Never test on training data
  2. Walk-forward analysis: Simulate production deployment
  3. Stress testing: How does the model perform in extreme scenarios?
  4. Benchmark comparison: Does ML beat simple baselines?

Example from renewable portfolio optimization:

  • Training period: 2022-2023 (build models and optimal allocation)
  • Test period: 2024 (validate)
  • ML models properly validated on holdout data before deployment

5. Production ML ≠ Academic ML

What changes in production:

  • Error handling: Models must fail gracefully
  • Logging: Track predictions, errors, data quality issues
  • Monitoring: Detect model degradation over time
  • Retraining: Automate model updates as new data arrives
  • Interpretability: Explain predictions to non-technical stakeholders

My Nord Pool surveillance system implements production-grade ML: automated anomaly detection, severity classification, comprehensive logging, and regulatory compliance checks. The code is engineered for reliability, not just accuracy.

Reflections on Revisiting Fundamentals

After building several production ML systems—energy trading strategies, portfolio optimization models, market surveillance tools—I’ve come to appreciate that fundamentals matter more over time, not less.

Each new project reveals deeper connections:

  • Random Forest ensemble methods parallel portfolio diversification strategies
  • Classification algorithms share optimization techniques with regression
  • Feature engineering borrows from domain expertise in energy markets

The algorithms I learned years ago—linear regression, logistic regression, random forests—still power most of my production systems. Deep learning gets the headlines, but classical ML combined with domain expertise often delivers better results with less complexity.

What I’d like to exploring next:

  • Gradient boosting (XGBoost, LightGBM) for tabular data – potentially better than Random Forest for my trading signals
  • Time series ML (LSTM, transformers) for electricity price forecasting
  • Bayesian optimization for hyperparameter tuning
  • Causal inference to distinguish correlation from causation in trading signals

The beauty of machine learning? The fundamentals remain constant, but the applications are endless.


What are you working on? What ML concepts do you wish you’d understood better from the start? Let’s discuss in the comments!

Leave a Reply

Trending

Discover more from Convergence Point

Subscribe now to keep reading and get access to the full archive.

Continue reading