Neural Network Toolbox

A Field Guide to Embedded Neural Network Algorithms

Gabriel Francisco dos Santos

gabriel.fra.santos@gmail.com

October 01, 2026

1 Activation Functions

1.1 Activation Functions

1.1.1 Overview & Motivation

An activation function \(f\) is the non-linear transformation applied after the affine map in each neural network layer:

\[a = f(z) = f(W x + b)\]

Without activation functions, stacking layers would collapse into a single affine transformation — the network could only represent linear mappings regardless of depth. Activation functions are what give neural networks their expressive power.

The choice of activation function controls gradient flow during back-propagation, output range, and computational cost — all critical on resource-constrained embedded targets.

1.1.2 Mathematical Theory

Activation functions

1.1.2.1 Forward and Backward

Every activation function provides two operations on a pre-activation vector \(z \in \mathbb{R}^k\) with output \(a = f(z)\):

Operation Definition Purpose
Forward \(a = f(z)\) Transform the pre-activation
Backward \(\delta = J_f(z)^T g\), with \(g = \partial \mathcal{L} / \partial a\) Propagate the upstream gradient to \(z\)

\(J_f(z) = \partial a / \partial z\) is the \(k \times k\) Jacobian. For the element-wise functions (ReLU, Leaky ReLU, Sigmoid, Tanh), \(a_i\) depends only on \(z_i\), so the Jacobian is diagonal and the backward pass reduces to a Hadamard product:

\[\delta = g \odot f'(z), \qquad \delta_i = g_i \, f'(z_i)\]

For Softmax, every output depends on every input, so the Jacobian is dense; the backward pass is still computed in \(O(k)\) as a Jacobian-vector product (see Softmax). The derivatives of Sigmoid and Tanh are written in terms of the output \(a\), so the backward pass reuses the cached forward result instead of re-evaluating the exponential.

1.1.2.2 Catalogue

1.1.2.2.1 ReLU

\[f(z) = \max(0, z), \qquad f'(z) = \begin{cases} 1 & z > 0 \\ 0 & z \le 0 \end{cases}\]

1.1.2.2.2 Leaky ReLU

\[f(z) = \begin{cases} z & z > 0 \\ \alpha z & z \le 0 \end{cases}, \qquad f'(z) = \begin{cases} 1 & z > 0 \\ \alpha & z \le 0 \end{cases}\]

where \(\alpha \in [0, 1)\) is a small constant (default \(0.01\)). Prevents dead neurons by allowing a small gradient for \(z < 0\); \(\alpha = 0\) reduces to ReLU.

1.1.2.2.3 Sigmoid

\[f(z) = \frac{1}{1 + e^{-z}}, \qquad f'(z) = f(z)\,\bigl(1 - f(z)\bigr)\]

The naive formula overflows \(e^{-z}\) for large negative \(z\). The numerically stable evaluation branches on the sign so the exponent is never positive:

\[f(z) = \begin{cases} \dfrac{1}{1 + e^{-z}} & z \ge 0 \\[2ex] \dfrac{e^{z}}{1 + e^{z}} & z < 0 \end{cases}\]

Both branches are the same function; each only evaluates \(e^{-|z|} \in (0, 1]\).

1.1.2.2.4 Tanh

\[f(z) = \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}}, \qquad f'(z) = 1 - f(z)^2\]

1.1.2.2.5 Softmax

For a vector \(z \in \mathbb{R}^k\):

\[a_i = f(z)_i = \frac{e^{z_i}}{\sum_{j=1}^{k} e^{z_j}}\]

Jacobian. Differentiating the quotient gives

\[\frac{\partial a_i}{\partial z_j} = a_i\,(\mathbb{1}[i = j] - a_j), \qquad J = \operatorname{diag}(a) - a\,a^T\]

Backward as a Jacobian-vector product. \(J\) is symmetric, so

\[\delta = J^T g = a \odot g - a\,(a^T g), \qquad \delta_i = a_i \Bigl(g_i - \sum_{k} g_k a_k\Bigr)\]

One dot product \(\langle g, a \rangle\) and one element-wise pass: \(O(k)\) time and \(O(1)\) extra space. The \(k \times k\) Jacobian is never formed.

1.1.3 Complexity Analysis

Activation Forward (per vector of \(k\)) Backward (per vector of \(k\)) Notes
ReLU \(O(k)\) — one comparison each \(O(k)\) — comparison + multiply Fastest
Leaky ReLU \(O(k)\) — comparison + multiply \(O(k)\) Negligible overhead vs ReLU
Sigmoid \(O(k)\) — one \(\exp\) each \(O(k)\) — reuses the output: \(a(1 - a)\) One transcendental per element
Tanh \(O(k)\) — one \(\tanh\) each \(O(k)\) — reuses the output: \(1 - a^2\) One transcendental per element
Softmax \(O(k)\) — max, \(k\) exps, one sum \(O(k)\) — Jacobian-vector product \(a \odot (g - \langle g, a \rangle)\) Full Jacobian never formed

All activations need \(O(1)\) extra space beyond the input, output and gradient vectors the layer already holds.

1.1.4 Step-by-Step Walkthrough

Element-wise case. ReLU on the pre-activations \(z = [-0.5, \; 1.2, \; 0.0]\).

Forward:

Neuron \(z\) \(\text{ReLU}(z)\)
1 \(-0.5\) \(0.0\)
2 \(1.2\) \(1.2\)
3 \(0.0\) \(0.0\)

Backward with incoming gradient \(g = \partial \mathcal{L} / \partial a = [0.3, \; -0.7, \; 0.1]\):

Neuron \(f'(z)\) \(\delta = g \odot f'(z)\)
1 \(0\) (dead) \(0.3 \times 0 = 0\)
2 \(1\) \(-0.7 \times 1 = -0.7\)
3 \(0\) (at boundary) \(0.1 \times 0 = 0\)

Neuron 1 is dead — its gradient is zero and its weights will not update. If this persists across all training samples, the neuron is permanently inactive.

Vector case. Softmax on \(z = [1.0, \; 2.0, \; 0.5]\) with upstream gradient \(g = [0.3, \; -0.2, \; 0.1]\).

Step Computation Result
Shift by \(z_{\max} = 2.0\) \(z - z_{\max}\) \([-1.0, \; 0.0, \; -1.5]\)
Exponentials \(e^{z - z_{\max}}\) \([0.3679, \; 1.0, \; 0.2231]\)
Normalise (sum \(1.5910\)) \(a = e^{z - z_{\max}} / 1.5910\) \([0.2312, \; 0.6285, \; 0.1402]\)
Dot product \(\langle g, a \rangle\) \(-0.0423\)
Backward \(\delta_i = a_i (g_i - \langle g, a \rangle)\) \([0.0792, \; -0.0991, \; 0.0200]\)

The outputs sum to \(1\) and the gradient sums to \(0\) (\(\sum_i \delta_i = \langle g, a \rangle - \langle g, a \rangle \sum_i a_i = 0\)), reflecting the shift invariance.

1.1.5 Pitfalls & Edge Cases

1.1.6 Variants & Generalizations

Variant Key Difference
Identity (linear) \(f(z) = z\), \(f' = 1\); raw regression outputs and logits
PReLU (Parametric ReLU) \(\alpha\) is a learnable parameter per channel
ELU (Exponential LU) \(\alpha(e^z - 1)\) for \(z < 0\); smooth and zero-centered
GELU (Gaussian Error LU) \(z \cdot \Phi(z)\); used in Transformers
Swish / SiLU \(z \cdot \sigma(z)\); smooth, non-monotonic
Hard Sigmoid / Hard Tanh Piece-wise linear approximations; no transcendentals

1.1.7 Applications

1.1.8 Connections to Other Algorithms

graph TD
    Act["Activation Functions"]
    Layer["Dense Layer"]
    NN["Neural Network"]
    Loss["Loss Functions"]
    Act --> Layer
    Layer --> NN
    NN --> Loss
Component Relationship
Dense Layer Applies the activation after the affine transformation and calls its backward pass
Neural Network Activations enable the non-linear function approximation that makes deep networks useful
Loss Functions Sigmoid output feeds binary cross-entropy; categorical cross-entropy takes logits and applies softmax itself — no Softmax layer

2 Layers

2.1 Dense Layer

2.1.1 Overview & Motivation

A dense (fully-connected) layer is the most fundamental building block of a neural network. It maps an input vector \(a_{\text{in}} \in \mathbb{R}^n\) to an output vector \(a_{\text{out}} \in \mathbb{R}^m\) through a learnable affine transformation followed by an activation function:

\[a_{\text{out}} = f(W \, a_{\text{in}} + b)\]

Every input neuron is connected to every output neuron — hence “fully connected.” The layer’s parameters are the weight matrix \(W\) and bias vector \(b\); training adjusts these to minimize the loss.

In this library, input size, output size, and parameter count are all compile-time constants, enabling static allocation and dimension checking with zero runtime overhead. The caller supplies the initial weights; the biases start at zero.

2.1.2 Mathematical Theory

Dense layer forward and backward pass

2.1.2.1 Forward Pass

\[z = W \, a_{\text{in}} + b, \qquad a_{\text{out}} = f(z)\]

where \(W \in \mathbb{R}^{m \times n}\), \(b \in \mathbb{R}^m\), and \(f\) is the activation function. The layer caches \(a_{\text{in}}\), \(z\) and \(a_{\text{out}}\) for the backward pass.

2.1.2.2 Backward Pass

Given the upstream gradient \(g = \partial \mathcal{L} / \partial a_{\text{out}} \in \mathbb{R}^m\) for one sample:

  1. Pre-activation gradient — the transposed activation Jacobian applied to \(g\): \[\delta = \frac{\partial \mathcal{L}}{\partial z} = J_f(z)^T g\] For an element-wise activation the Jacobian is diagonal and this reduces to \(\delta = g \odot f'(z)\). For Softmax it is \(\delta_i = a_i \bigl(g_i - \sum_k g_k a_k\bigr)\) with \(a = a_{\text{out}}\), still \(O(m)\) (see Softmax).

  2. Weight gradient: \[\frac{\partial \mathcal{L}}{\partial W} = \delta \, a_{\text{in}}^T\]

  3. Bias gradient: \[\frac{\partial \mathcal{L}}{\partial b} = \delta\]

  4. Input gradient (propagated to the previous layer): \[\frac{\partial \mathcal{L}}{\partial a_{\text{in}}} = W^T \delta\]

The backward pass requires a preceding forward pass (it uses the cached \(a_{\text{in}}\), \(z\), \(a_{\text{out}}\)). Each call overwrites the stored parameter gradients with those of the current sample; they are not accumulated and are not yet exposed to an optimizer (roadmap N8).

2.1.2.3 Parameter Vector

The parameters are stored as one flat vector: \(W\) in row-major order (row \(i\) holds the weights of output neuron \(i\)) followed by \(b\):

\[\theta = [\, W_{1,1}, \ldots, W_{1,n}, \; W_{2,1}, \ldots, W_{m,n}, \; b_1, \ldots, b_m \,], \qquad P = m\,n + m = m(n + 1)\]

For a layer with 128 inputs and 64 outputs: \(P = 64 \times 129 = 8{,}256\) parameters.

2.1.3 Complexity Analysis

Operation Time Extra space
Forward (\(W a + b\), then \(f\)) \(O(m \cdot n)\) caches \(a_{\text{in}}\) (\(n\)), \(z\) (\(m\)), \(a_{\text{out}}\) (\(m\))
Backward (\(\delta\), \(\nabla W\), \(\nabla b\), \(W^T\delta\)) \(O(m \cdot n)\) \(O(m)\) temporary \(\delta\)

The matrix-vector products dominate both passes. For embedded networks (e.g. \(n = 32, m = 16\)), a single forward pass takes 512 multiply-accumulate operations.

Resident memory per layer. The object holds the parameters (\(P\) values), a parameter-gradient buffer of the same size (\(P\)), and the cached input, pre-activation, output and input-gradient vectors (\(2n + 2m\)):

\[\text{values} = 2P + 2n + 2m = 2mn + 4m + 2n\]

That is roughly twice the weight matrix. A \(256 \times 256\) float layer holds \(2 \times 65{,}792 + 1{,}024 = 132{,}608\) values, about \(518\,\text{KiB}\) — not the \(256\,\text{KiB}\) of the weights alone. The model adds a transient \(P\)-sized copy when parameters are exported or imported.

2.1.4 Step-by-Step Walkthrough

Layer: 3 inputs → 2 outputs, ReLU activation.

\[W = \begin{bmatrix} 0.5 & -0.3 & 0.8 \\ 0.1 & 0.7 & -0.2 \end{bmatrix}, \quad b = \begin{bmatrix} 0.1 \\ -0.1 \end{bmatrix}, \quad a_{\text{in}} = \begin{bmatrix} 1.0 \\ 0.5 \\ -1.0 \end{bmatrix}\]

Forward:

Step Computation Result
\(z = W a_{\text{in}} + b\) \([0.5 - 0.15 - 0.8 + 0.1,\; 0.1 + 0.35 + 0.2 - 0.1]\) \([-0.35,\; 0.55]^T\)
\(a_{\text{out}} = \text{ReLU}(z)\) \([\max(0, -0.35),\; \max(0, 0.55)]\) \([0.0,\; 0.55]^T\)

Backward with \(g = \partial \mathcal{L} / \partial a_{\text{out}} = [0.2,\; -0.4]^T\):

Step Computation Result
\(\delta = g \odot \text{ReLU}'(z)\) \([0.2 \cdot 0,\; -0.4 \cdot 1]\) \([0,\; -0.4]^T\)
\(\nabla W = \delta \, a_{\text{in}}^T\) row 1: all zeros; row 2: \(-0.4 \times [1, 0.5, -1]\) \(\begin{bmatrix}0 & 0 & 0\\-0.4 & -0.2 & 0.4\end{bmatrix}\)
\(\nabla b = \delta\) — \([0,\; -0.4]^T\)
\(\nabla a_{\text{in}} = W^T \delta\) \(-0.4 \times [0.1,\; 0.7,\; -0.2]\) \([-0.04,\; -0.28,\; 0.08]^T\)

In the flat parameter layout, \(\nabla\theta = [0, 0, 0, -0.4, -0.2, 0.4, 0, -0.4]\).

2.1.5 Pitfalls & Edge Cases

2.1.6 Variants & Generalizations

Variant Key Difference
Convolutional layer Weight sharing across spatial positions; \(O(k^2 \cdot c)\) parameters per filter instead of \(O(n \cdot m)\)
Recurrent layer Shares weights across time steps; adds a hidden state feedback connection
Batch normalization layer Normalizes activations to zero mean and unit variance; accelerates training
Dropout layer Randomly zeros activations during training; regularization effect
Sparse layer Only a subset of connections exist; reduces parameter count and computation

2.1.7 Applications

2.1.8 Connections to Other Algorithms

graph TD
    Layer["Dense Layer"]
    Act["Activation Functions"]
    Model["Model"]
    Opt["Optimizer"]
    LR["Linear Regression"]

    Act --> Layer
    Layer --> Model
    Model --> Opt
    Layer -.->|"identity activation, MSE loss"| LR
Component Relationship
Activation Functions Applied after the affine transformation; its Jacobian-vector product gives \(\delta\)
Model Chains multiple dense layers and concatenates their flat parameter vectors
Optimizer (numerical-toolbox-cpp) Updates the model’s flat parameter vector; feeding it the layer’s own gradients is roadmap N8/N16
Linear Regression (numerical-toolbox-cpp) A dense layer with an identity activation \(f(z) = z\) and MSE loss is linear regression; the identity activation is roadmap N1

3 Loss Functions

3.1 Loss Functions

3.1.1 Overview & Motivation

A loss function \(\mathcal{L}(\hat{y}, y)\) quantifies how far a model’s prediction \(\hat{y}\) is from the true target \(y\). Training a neural network means finding parameters \(\theta\) that minimize the expected loss over the training data:

\[\theta^* = \arg\min_\theta \; \mathbb{E}[\mathcal{L}(f_\theta(x), y)]\]

The loss function defines the entire learning objective — different losses lead to different optimal models even on the same data. It must also provide the gradient \(\nabla_{\hat{y}} \mathcal{L}\) that seeds back-propagation.

In this library every loss is an objective over one vector argument \(v \in \mathbb{R}^N\) — the \(N\) outputs of one sample — compared against a target \(y \in \mathbb{R}^N\) fixed when the loss is created. A regularization term \(R\) from numerical-toolbox-cpp is added to the same argument:

\[J(v) = \mathcal{L}(v, y) + R(v), \qquad \nabla J(v) = \nabla_v \mathcal{L}(v, y) + \nabla R(v)\]

For MSE, MAE and BCE the argument is the prediction \(v = \hat{y}\); for categorical cross-entropy it is the vector of logits \(v = z\).

3.1.2 Mathematical Theory

Loss functions

All data terms are means over the \(N\) outputs, so their gradients carry a \(1/N\) factor.

3.1.2.1 Mean Squared Error

\[\mathcal{L}_{\text{MSE}} = \frac{1}{N} \sum_{i=1}^{N} (\hat{y}_i - y_i)^2, \qquad \frac{\partial \mathcal{L}}{\partial \hat{y}_i} = \frac{2}{N}(\hat{y}_i - y_i)\]

3.1.2.2 Mean Absolute Error

\[\mathcal{L}_{\text{MAE}} = \frac{1}{N} \sum_{i=1}^{N} |\hat{y}_i - y_i|, \qquad \frac{\partial \mathcal{L}}{\partial \hat{y}_i} = \frac{1}{N} \operatorname{sign}(\hat{y}_i - y_i)\]

3.1.2.3 Binary Cross-Entropy

The input is a vector of probabilities \(p = \hat{y} \in [0, 1]^N\) (e.g. a Sigmoid output) and \(y_i \in [0, 1]\). Each probability is first clamped to \(\tilde{p}_i = \operatorname{clamp}(p_i, \varepsilon, 1 - \varepsilon)\) with \(\varepsilon = 10^{-7}\):

\[\mathcal{L}_{\text{BCE}} = -\frac{1}{N}\sum_{i=1}^{N} \left[ y_i \log \tilde{p}_i + (1 - y_i) \log(1 - \tilde{p}_i) \right]\]

\[\frac{\partial \mathcal{L}}{\partial p_i} = \frac{1}{N} \cdot \frac{\tilde{p}_i - y_i}{\tilde{p}_i(1 - \tilde{p}_i)}\]

3.1.2.4 Categorical Cross-Entropy

The input is a vector of logits \(z \in \mathbb{R}^k\) (\(N = k\) classes), not probabilities. The loss fuses softmax and cross-entropy:

\[\mathcal{L}_{\text{CCE}} = -\sum_{i=1}^{k} y_i \log \operatorname{softmax}(z)_i = \Bigl(\sum_{i} y_i\Bigr) \operatorname{LSE}(z) - \sum_{i} y_i z_i, \qquad \operatorname{LSE}(z) = \log \sum_{j} e^{z_j}\]

\[\frac{\partial \mathcal{L}}{\partial z_i} = \operatorname{softmax}(z)_i \sum_{j} y_j - y_i\]

3.1.3 Complexity Analysis

Loss Cost Gradient Transcendentals
MSE \(O(N)\) \(O(N)\) none
MAE \(O(N)\) \(O(N)\) none (sign)
BCE \(O(N)\) \(O(N)\) \(2N\) logarithms (cost only)
CCE \(O(k)\) \(O(k)\) \(k\) exponentials + 1 logarithm per call

All losses are linear in the output dimension and use \(O(1)\) extra space besides the returned gradient; the regularization term adds its own \(O(N)\). The cost is negligible compared to the dense-layer matrix products.

3.1.4 Step-by-Step Walkthrough

Categorical cross-entropy on logits. 3 classes, target \(y = [0, 1, 0]\) (class 2), logits \(z = [1.0, 2.0, 0.5]\).

Step Computation Result
Max shift \(z_{\max} = 2.0\), \(z - z_{\max}\) \([-1.0,\; 0.0,\; -1.5]\)
Shifted exponentials \(e^{z - z_{\max}}\) \([0.3679,\; 1.0,\; 0.2231]\), sum \(1.5910\)
Log-sum-exp \(2.0 + \log 1.5910\) \(2.4644\)
Cost \(\operatorname{LSE}(z) \cdot 1 - z_2\) \(2.4644 - 2.0 = 0.4644\)
Softmax \(e^{z_i - \operatorname{LSE}(z)}\) \([0.2312,\; 0.6285,\; 0.1402]\)
Gradient \(\operatorname{softmax}(z) - y\) \([0.2312,\; -0.3715,\; 0.1402]\)

The cost equals \(-\log 0.6285\), the negative log-probability of the true class. The gradient pushes the true logit up and the others down, and sums to zero.

For comparison — MSE on the same probabilities \(\hat{y} = [0.2312, 0.6285, 0.1402]\):

\[\mathcal{L}_{\text{MSE}} = \frac{1}{3}\bigl[0.2312^2 + 0.3715^2 + 0.1402^2\bigr] = 0.0704, \qquad \nabla = \frac{2}{3}(\hat{y} - y) = [0.1541,\; -0.2476,\; 0.0935]\]

Binary cross-entropy with \(p = [0.9, 0.2]\) and \(y = [1, 0]\) (\(N = 2\), no clamping active):

\[\mathcal{L}_{\text{BCE}} = -\tfrac{1}{2}\bigl[\log 0.9 + \log 0.8\bigr] = 0.1643, \qquad \nabla = \tfrac{1}{2}\Bigl[\tfrac{0.9 - 1}{0.9 \cdot 0.1},\; \tfrac{0.2 - 0}{0.2 \cdot 0.8}\Bigr] = [-0.5556,\; 0.625]\]

3.1.5 Pitfalls & Edge Cases

3.1.6 Variants & Generalizations

Variant Key Difference
Huber loss Quadratic for small errors, linear for large; robust regression
Logit-input BCE Fuses Sigmoid and BCE: \(\operatorname{softplus}(z) - yz\); no clamp needed
Focal loss Down-weights well-classified examples; addresses class imbalance
KL divergence Measures distance between two distributions; used in variational inference
Hinge loss Margin-based; used in SVMs and some neural classifiers
Contrastive loss Learns similarity metrics; used in Siamese networks

3.1.7 Applications

3.1.8 Connections to Other Algorithms

graph TD
    Loss["Loss Functions"]
    Act["Activation Functions"]
    Opt["Optimizer"]
    Reg["Regularization"]
    Model["Model"]
    LR["Linear Regression"]

    Act -->|"output activation must match loss"| Loss
    Loss --> Opt
    Reg -->|"added to loss"| Loss
    Loss --> Model
    Loss -.->|"MSE + normal equation"| LR
Component Relationship
Activation Functions Output activation must match: Sigmoid ↔︎ BCE, identity (logits) ↔︎ CCE, identity ↔︎ MSE/MAE
Model The model’s training step minimises an objective of this form over the flat parameter vector \(\theta\)
Optimizer (numerical-toolbox-cpp) Minimises \(J\) using its cost and gradient
Regularization (numerical-toolbox-cpp) Adds \(R(v)\) and \(\nabla R(v)\) to the data term
Linear Regression (numerical-toolbox-cpp) Solved analytically when the loss is MSE and the model is linear

4 Model

4.1 Neural Networks

4.1.1 Overview & Motivation

A neural network is a parameterized function \(f_\theta: \mathbb{R}^n \to \mathbb{R}^m\) built by composing simple, differentiable transformations called layers. Each layer applies an affine map followed by a non-linear activation, and the whole composition is trained by gradient-based optimization.

Feed-forward network: input layer, hidden layers, output layer

Neural networks are powerful because of the universal approximation theorem: a single hidden layer with enough neurons can approximate any continuous function on a compact set to arbitrary accuracy. In practice, depth (many layers) is more parameter-efficient than width for learning hierarchical features.

This library provides a minimal, statically-sized neural network framework designed for embedded inference — no heap allocation, no dynamic shapes, full compile-time dimension checking. The forward and backward passes described below are implemented; the parameter update loop that turns them into on-device training is roadmap N8/N9/N16 (see Model).

4.1.2 Mathematical Theory

4.1.2.1 Forward Propagation

Given \(L\) layers, the network computes:

\[a_0 = x\] \[z_\ell = W_\ell \, a_{\ell-1} + b_\ell, \quad \ell = 1, \ldots, L\] \[a_\ell = f_\ell(z_\ell)\] \[\hat{y} = a_L\]

where \(W_\ell \in \mathbb{R}^{n_\ell \times n_{\ell-1}}\) are weights, \(b_\ell \in \mathbb{R}^{n_\ell}\) are biases, and \(f_\ell\) is the activation function for layer \(\ell\).

4.1.2.2 Loss Function

Training minimizes a scalar loss \(\mathcal{L}(\hat{y}, y)\) that measures how far the prediction \(\hat{y}\) is from the target \(y\). Common choices (see Loss Functions):

Loss Formula Input Use Case
MSE \(\frac{1}{N}\sum_i(\hat{y}_i - y_i)^2\) predictions Regression
BCE \(-\frac{1}{N}\sum_i[y_i \log p_i + (1-y_i)\log(1-p_i)]\) probabilities \(p\) Binary classification
CCE \(-\sum_i y_i \log \operatorname{softmax}(z)_i\) logits \(z\) Multi-class classification

4.1.2.3 Backpropagation

Backpropagation computes \(\nabla_\theta \mathcal{L}\) via the chain rule, working from the output layer backward. With \(g_L = \nabla_{a_L} \mathcal{L}\) and \(J_{f_\ell}\) the Jacobian of the activation:

\[\delta_\ell = J_{f_\ell}(z_\ell)^T g_\ell, \qquad g_{\ell-1} = W_\ell^T \delta_\ell\]

For element-wise activations \(J_{f_\ell}\) is diagonal and \(\delta_\ell = g_\ell \odot f_\ell'(z_\ell)\); for Softmax it is the \(O(n_\ell)\) product \(\delta_\ell = a_\ell \odot (g_\ell - \langle g_\ell, a_\ell \rangle)\) (see Activation Functions). The gradients with respect to the parameters are:

\[\frac{\partial \mathcal{L}}{\partial W_\ell} = \delta_\ell \, a_{\ell-1}^T, \qquad \frac{\partial \mathcal{L}}{\partial b_\ell} = \delta_\ell\]

4.1.2.4 Parameter Update

An optimizer uses the gradients to update the parameter vector \(\theta\):

\[\theta_{t+1} = \theta_t - \eta \, \nabla_\theta \mathcal{L}\]

where \(\eta\) is the learning rate. More sophisticated optimizers (momentum, Adam) modify this basic rule. With a mini-batch of \(B\) samples, \(\nabla_\theta \mathcal{L}\) is the average of the per-sample gradients.

4.1.3 Complexity Analysis

Phase Time Space
Forward pass \(O\!\left(\sum_{\ell=1}^L n_\ell \cdot n_{\ell-1}\right)\) \(O\!\left(\sum_\ell n_\ell\right)\) activations
Backward pass Same as forward Same + \(O(P)\) gradient storage
Parameter update \(O(P)\) \(O(P)\) optimizer state

where \(P = \sum_\ell (n_\ell \cdot n_{\ell-1} + n_\ell)\) is the total parameter count. Both passes cost about one multiply-accumulate per weight, so a network with \(P = 10{,}000\) parameters needs on the order of \(10^4\) MACs per sample and pass.

4.1.4 Step-by-Step Walkthrough

Network: 2 inputs → 2 hidden (ReLU) → 1 output (Sigmoid), trained with binary cross-entropy on one XOR sample.

Architecture:

graph LR
    x1((x₁)) --> h1((h₁))
    x1 --> h2((h₂))
    x2((x₂)) --> h1
    x2 --> h2
    h1 --> y((ŷ))
    h2 --> y

Initial parameters:

\[W_1 = \begin{bmatrix} 0.2 & 0.4 \\ -0.3 & 0.1 \end{bmatrix}, \quad b_1 = \begin{bmatrix} 0.1 \\ 0.2 \end{bmatrix}, \quad W_2 = \begin{bmatrix} 0.5 & -0.4 \end{bmatrix}, \quad b_2 = [0]\]

Forward pass with input \(x = [1, 0]^T\), target \(y = 1\):

Step Computation Result
Hidden pre-activation \(z_1 = W_1 x + b_1\) \([0.3, -0.1]^T\)
Hidden activation \(a_1 = \text{ReLU}(z_1)\) \([0.3, 0.0]^T\)
Output pre-activation \(z_2 = W_2 a_1 + b_2\) \(0.15\)
Output activation \(\hat{y} = \sigma(z_2)\) \(0.5374\)
Loss \(\mathcal{L} = -(y\log\hat{y} + (1-y)\log(1-\hat{y}))\) \(0.6210\)

Backward pass (Sigmoid followed by BCE simplifies to \(\delta_2 = \hat{y} - y\)):

Step Computation Result
Output gradient \(\delta_2 = \hat{y} - y\) \(-0.4626\)
\(\nabla W_2\) \(\delta_2 \, a_1^T\) \([-0.1388, 0]\)
\(\nabla b_2\) \(\delta_2\) \(-0.4626\)
Hidden gradient \(\delta_1 = (W_2^T \delta_2) \odot \text{ReLU}'(z_1)\) \([-0.2313, 0]^T\)
\(\nabla W_1\) \(\delta_1 \, x^T\) \([[-0.2313, 0], [0, 0]]\)
\(\nabla b_1\) \(\delta_1\) \([-0.2313, 0]^T\)

Update with \(\eta = 0.1\): \(W_2 \leftarrow [0.5139, -0.4]\), \(b_2 \leftarrow 0.0463\), \(W_{1,11} \leftarrow 0.2231\), \(b_{1,1} \leftarrow 0.1231\) (all other parameters have zero gradient). Repeating the forward pass gives \(\hat{y} = \sigma(0.2242) = 0.5558\) and \(\mathcal{L} = 0.5873\) — the loss decreased. Training repeats this step over all four XOR samples.

4.1.5 Pitfalls & Edge Cases

4.1.6 Variants & Generalizations

Variant Key Difference
Convolutional Neural Network (CNN) Layers share weights spatially; efficient for image/signal data
Recurrent Neural Network (RNN) Layers share weights across time steps; models sequences
Residual Network (ResNet) Skip connections mitigate vanishing gradients in very deep networks
Transformer Attention-based; no recurrence; state-of-the-art for sequences
Quantized Neural Network Weights and activations in low-bit integers; smaller and faster on MCUs

4.1.7 Applications

4.1.8 Connections to Other Algorithms

graph TD
    NN["Neural Network"]
    Layer["Dense Layer"]
    Act["Activation Functions"]
    Loss["Loss Functions"]
    Opt["Optimizer"]
    Reg["Regularization"]
    Model["Model"]
    LR["Linear Regression"]

    Layer --> NN
    Act --> NN
    Loss --> NN
    Opt --> NN
    Reg --> NN
    NN --> Model
    NN -.->|"single layer, identity activation, MSE loss"| LR
Component Relationship
Dense Layer The fundamental building block; computes affine transformations
Activation Functions Introduce non-linearity after each layer
Loss Functions Define the training objective
Optimizer (numerical-toolbox-cpp) Drives parameter updates via gradient descent
Regularization (numerical-toolbox-cpp) Penalizes complexity to prevent overfitting
Model Composes layers into one network with a flat parameter vector
Linear Regression (numerical-toolbox-cpp) Special case: single layer, identity activation, MSE loss

4.2 Model (Sequential Composition)

4.2.1 Overview & Motivation

A model composes a sequence of layers — today dense layers — into a single function \(f_\theta: \mathbb{R}^n \to \mathbb{R}^m\). It:

  1. Chains layers so the output of each feeds into the next (forward pass).
  2. Propagates the output gradient backward through the chain (backward pass), returning \(\partial \mathcal{L} / \partial x\); each layer computes its own parameter gradients on the way.
  3. Exports and imports all layer parameters as one flat vector \(\theta \in \mathbb{R}^P\).
  4. Verifies dimensional compatibility at compile time: the model input size, every layer-to-layer interface and the model output size must agree, and there must be at least one layer.
  5. Offers a parameter-space training step: an optimizer minimises a user-supplied objective \(J(\theta)\) over the flat parameter vector and the minimiser is loaded into the layers.

The model is fully statically typed — layer dimensions, parameter counts, and memory footprints are all known at compile time, enabling zero-overhead abstraction on embedded targets.

4.2.2 Mathematical Theory

Model composition, parameter vector and training step

4.2.2.1 Composition

For \(L\) layers with transformations \(f_1, f_2, \ldots, f_L\):

\[\hat{y} = (f_L \circ f_{L-1} \circ \cdots \circ f_1)(x) = f_L(f_{L-1}(\ldots f_1(x) \ldots))\]

Each \(f_\ell\) is a dense layer: \(a_\ell = f_\ell(a_{\ell-1}) = \sigma_\ell(W_\ell a_{\ell-1} + b_\ell)\) with \(a_0 = x\), \(W_\ell \in \mathbb{R}^{m_\ell \times n_\ell}\) and \(n_{\ell+1} = m_\ell\).

4.2.2.2 Parameter Vector

All weights and biases are concatenated in layer order; each \(W_\ell\) is flattened row-major (row \(i\) = weights of output neuron \(i\)), followed by its bias:

\[\theta = [\, \operatorname{rows}(W_1),\ b_1,\ \operatorname{rows}(W_2),\ b_2,\ \ldots,\ \operatorname{rows}(W_L),\ b_L \,] \in \mathbb{R}^P, \qquad P = \sum_{\ell=1}^L m_\ell(n_\ell + 1)\]

Export and import use exactly this order, so exporting and re-importing is the identity.

4.2.2.3 Forward Pass (Chained Evaluation)

graph LR
    X["x ∈ ℝⁿ"] --> L1["Layer 1"] --> L2["Layer 2"] --> Ldots["⋯"] --> LL["Layer L"] --> Y["ŷ ∈ ℝᵐ"]

4.2.2.4 Backward Pass (Reverse Chain Rule)

Given \(g_L = \partial \mathcal{L} / \partial \hat{y}\), the layers are visited in reverse. Layer \(\ell\) receives \(g_\ell = \partial \mathcal{L} / \partial a_\ell\) and computes

\[\delta_\ell = J_{\sigma_\ell}(z_\ell)^T g_\ell, \qquad \nabla W_\ell = \delta_\ell\, a_{\ell-1}^T, \qquad \nabla b_\ell = \delta_\ell, \qquad g_{\ell-1} = W_\ell^T \delta_\ell\]

so each layer applies its own activation derivative before handing \(g_{\ell-1}\) to its predecessor. The model returns \(g_0 = \partial \mathcal{L} / \partial x\). The per-layer parameter gradients stay inside the layers; exporting them as a flat \(\nabla_\theta \mathcal{L}\) is roadmap N8.

4.2.2.5 Training Step (Parameter-Space Optimisation)

The training step takes an optimizer, an objective \(J: \mathbb{R}^P \to \mathbb{R}\) with its gradient \(\nabla J\), and a starting point \(\theta_0\):

\[\theta^* = \operatorname{Optimize}(J, \theta_0) \approx \arg\min_\theta J(\theta), \qquad \text{then load } \theta^* \text{ into the layers}\]

It does not iterate over a dataset and does not call the forward or backward pass: the objective must itself encode whatever it measures about \(\theta\) (for example a user-written empirical risk). If a per-sample loss is passed as \(J\), its argument is \(\theta\) and its target is a parameter vector — it regresses the parameters onto that target. End-to-end back-propagation training — per-sample loss on \(\hat{y}\) (N9), exported parameter gradients (N8) and a mini-batch gradient step (N16) — is on the roadmap.

graph LR
    J["Objective J(θ), ∇J(θ)"] --> OPT["Optimizer from θ₀"] --> TH["θ*"] --> SET["Load θ* into layers"]

4.2.3 Complexity Analysis

Operation Time Space
Forward pass \(O(P)\) per-layer caches \(O(\sum_\ell (n_\ell + m_\ell))\)
Backward pass \(O(P)\) per-layer gradient buffers \(O(P)\) (resident)
Parameter export \(O(P)\) \(O(P)\) flat vector returned
Parameter import \(O(P)\) \(O(\max_\ell P_\ell)\) transient per-layer copy
Training step optimizer-dependent optimizer state + \(O(P)\)

Forward and backward are linear in the total parameter count \(P\) (one multiply-accumulate per weight). The resident model is the sum of its layers — about \(2P\) values plus the activation caches for dense layers (see Dense Layer).

4.2.4 Step-by-Step Walkthrough

Model: 2 → 3 → 1. Layer 1 uses ReLU, layer 2 uses Sigmoid.

Compile-time verification chain:

Check Condition Status
At least one layer \(L = 2 \ge 1\) ✓
Layer 1 input size = model input size \(2 = 2\) ✓
Layer 1 output size = layer 2 input size \(3 = 3\) ✓
Layer 2 output size = model output size \(1 = 1\) ✓
Every layer shares the model’s value type float ✓

Parameters (\(P = 3(2 + 1) + 1(3 + 1) = 13\)):

\[W_1 = \begin{bmatrix} -0.5 & 0.2 \\ 0.4 & 0.6 \\ 0.1 & 0.2 \end{bmatrix}, \quad b_1 = \begin{bmatrix} 0.1 \\ 0.1 \\ 0.1 \end{bmatrix}, \quad W_2 = \begin{bmatrix} 0.3 & 0.5 & 0.2 \end{bmatrix}, \quad b_2 = [0.03]\]

Index Parameter Values
0–5 \(W_1\) row-major (\(3 \times 2\)) \(-0.5, 0.2, 0.4, 0.6, 0.1, 0.2\)
6–8 \(b_1\) \(0.1, 0.1, 0.1\)
9–11 \(W_2\) row-major (\(1 \times 3\)) \(0.3, 0.5, 0.2\)
12 \(b_2\) \(0.03\)

Forward pass with \(x = [1.0, 0.5]^T\):

Step Computation Result
Layer 1 \(z_1 = W_1 x + b_1\) \([-0.3,\; 0.8,\; 0.3]^T\)
\(a_1 = \text{ReLU}(z_1)\) \([0.0,\; 0.8,\; 0.3]^T\)
Layer 2 \(z_2 = W_2 a_1 + b_2 = 0.4 + 0.06 + 0.03\) \(0.49\)
\(\hat{y} = \sigma(z_2)\) \(0.6201\)

Backward pass with loss gradient \(g_2 = \partial \mathcal{L} / \partial \hat{y} = [0.12]\):

Step Computation Result
Layer 2 \(\delta_2 = g_2 \, \sigma'(z_2) = 0.12 \cdot 0.6201 \cdot 0.3799\) \(0.02827\)
\(\nabla W_2 = \delta_2 a_1^T\), \(\nabla b_2 = \delta_2\) \([0,\; 0.02262,\; 0.00848]\), \(0.02827\)
\(g_1 = W_2^T \delta_2\) (handed to layer 1) \([0.00848,\; 0.01413,\; 0.00565]^T\)
Layer 1 \(\delta_1 = g_1 \odot \text{ReLU}'(z_1)\) \([0,\; 0.01413,\; 0.00565]^T\)
\(\nabla W_1 = \delta_1 x^T\), \(\nabla b_1 = \delta_1\) rows \([0, 0]\), \([0.01413, 0.00707]\), \([0.00565, 0.00283]\)
\(g_0 = W_1^T \delta_1\) (returned by the model) \([0.00622,\; 0.00961]^T\)

Note where \(\text{ReLU}'(z_1)\) enters: inside layer 1’s own backward pass, not in layer 2’s.

4.2.5 Pitfalls & Edge Cases

4.2.6 Variants & Generalizations

Variant Key Difference
Sequential model (dynamic) Layers stored in a container; dimension checked at runtime instead of compile time
Functional API Supports branching and merging (DAG topology instead of linear chain)
Residual model Adds skip connections: \(a_{\ell+2} = f_{\ell+1}(a_\ell) + a_\ell\)
Recurrent model Unrolls the same layer across time steps

4.2.7 Applications

4.2.8 Connections to Other Algorithms

graph TD
    Model["Model"]
    Layer["Dense Layer"]
    Loss["Loss Functions"]
    Opt["Optimizer"]
    Reg["Regularization"]
    NN["Neural Network"]

    Layer --> Model
    Model --> Opt
    Loss --> Opt
    Reg --> Loss
    Model --> NN
Component Relationship
Dense Layer The model is a fixed sequence of layers evaluated in order
Neural Network Overview of forward propagation, back-propagation and the parameter update the model builds on
Loss Functions Objectives of the same form; the training step accepts one sized to the parameter vector
Optimizer (numerical-toolbox-cpp) Minimises the objective over the flat parameter vector and returns the minimiser
Regularization (numerical-toolbox-cpp) Added to the objective’s argument by the loss

5 References