A Field Guide to Embedded Neural Network Algorithms
October 01, 2026
An activation function \(f\) is the non-linear transformation applied after the affine map in each neural network layer:
\[a = f(z) = f(W x + b)\]
Without activation functions, stacking layers would collapse into a single affine transformation — the network could only represent linear mappings regardless of depth. Activation functions are what give neural networks their expressive power.
The choice of activation function controls gradient flow during back-propagation, output range, and computational cost — all critical on resource-constrained embedded targets.
Every activation function provides two operations on a pre-activation vector \(z \in \mathbb{R}^k\) with output \(a = f(z)\):
| Operation | Definition | Purpose |
|---|---|---|
| Forward | \(a = f(z)\) | Transform the pre-activation |
| Backward | \(\delta = J_f(z)^T g\), with \(g = \partial \mathcal{L} / \partial a\) | Propagate the upstream gradient to \(z\) |
\(J_f(z) = \partial a / \partial z\) is the \(k \times k\) Jacobian. For the element-wise functions (ReLU, Leaky ReLU, Sigmoid, Tanh), \(a_i\) depends only on \(z_i\), so the Jacobian is diagonal and the backward pass reduces to a Hadamard product:
\[\delta = g \odot f'(z), \qquad \delta_i = g_i \, f'(z_i)\]
For Softmax, every output depends on every input, so the Jacobian is dense; the backward pass is still computed in \(O(k)\) as a Jacobian-vector product (see Softmax). The derivatives of Sigmoid and Tanh are written in terms of the output \(a\), so the backward pass reuses the cached forward result instead of re-evaluating the exponential.
\[f(z) = \max(0, z), \qquad f'(z) = \begin{cases} 1 & z > 0 \\ 0 & z \le 0 \end{cases}\]
\[f(z) = \begin{cases} z & z > 0 \\ \alpha z & z \le 0 \end{cases}, \qquad f'(z) = \begin{cases} 1 & z > 0 \\ \alpha & z \le 0 \end{cases}\]
where \(\alpha \in [0, 1)\) is a small constant (default \(0.01\)). Prevents dead neurons by allowing a small gradient for \(z < 0\); \(\alpha = 0\) reduces to ReLU.
\[f(z) = \frac{1}{1 + e^{-z}}, \qquad f'(z) = f(z)\,\bigl(1 - f(z)\bigr)\]
The naive formula overflows \(e^{-z}\) for large negative \(z\). The numerically stable evaluation branches on the sign so the exponent is never positive:
\[f(z) = \begin{cases} \dfrac{1}{1 + e^{-z}} & z \ge 0 \\[2ex] \dfrac{e^{z}}{1 + e^{z}} & z < 0 \end{cases}\]
Both branches are the same function; each only evaluates \(e^{-|z|} \in (0, 1]\).
\[f(z) = \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}}, \qquad f'(z) = 1 - f(z)^2\]
For a vector \(z \in \mathbb{R}^k\):
\[a_i = f(z)_i = \frac{e^{z_i}}{\sum_{j=1}^{k} e^{z_j}}\]
Jacobian. Differentiating the quotient gives
\[\frac{\partial a_i}{\partial z_j} = a_i\,(\mathbb{1}[i = j] - a_j), \qquad J = \operatorname{diag}(a) - a\,a^T\]
Backward as a Jacobian-vector product. \(J\) is symmetric, so
\[\delta = J^T g = a \odot g - a\,(a^T g), \qquad \delta_i = a_i \Bigl(g_i - \sum_{k} g_k a_k\Bigr)\]
One dot product \(\langle g, a \rangle\) and one element-wise pass: \(O(k)\) time and \(O(1)\) extra space. The \(k \times k\) Jacobian is never formed.
| Activation | Forward (per vector of \(k\)) | Backward (per vector of \(k\)) | Notes |
|---|---|---|---|
| ReLU | \(O(k)\) — one comparison each | \(O(k)\) — comparison + multiply | Fastest |
| Leaky ReLU | \(O(k)\) — comparison + multiply | \(O(k)\) | Negligible overhead vs ReLU |
| Sigmoid | \(O(k)\) — one \(\exp\) each | \(O(k)\) — reuses the output: \(a(1 - a)\) | One transcendental per element |
| Tanh | \(O(k)\) — one \(\tanh\) each | \(O(k)\) — reuses the output: \(1 - a^2\) | One transcendental per element |
| Softmax | \(O(k)\) — max, \(k\) exps, one sum | \(O(k)\) — Jacobian-vector product \(a \odot (g - \langle g, a \rangle)\) | Full Jacobian never formed |
All activations need \(O(1)\) extra space beyond the input, output and gradient vectors the layer already holds.
Element-wise case. ReLU on the pre-activations \(z = [-0.5, \; 1.2, \; 0.0]\).
Forward:
| Neuron | \(z\) | \(\text{ReLU}(z)\) |
|---|---|---|
| 1 | \(-0.5\) | \(0.0\) |
| 2 | \(1.2\) | \(1.2\) |
| 3 | \(0.0\) | \(0.0\) |
Backward with incoming gradient \(g = \partial \mathcal{L} / \partial a = [0.3, \; -0.7, \; 0.1]\):
| Neuron | \(f'(z)\) | \(\delta = g \odot f'(z)\) |
|---|---|---|
| 1 | \(0\) (dead) | \(0.3 \times 0 = 0\) |
| 2 | \(1\) | \(-0.7 \times 1 = -0.7\) |
| 3 | \(0\) (at boundary) | \(0.1 \times 0 = 0\) |
Neuron 1 is dead — its gradient is zero and its weights will not update. If this persists across all training samples, the neuron is permanently inactive.
Vector case. Softmax on \(z = [1.0, \; 2.0, \; 0.5]\) with upstream gradient \(g = [0.3, \; -0.2, \; 0.1]\).
| Step | Computation | Result |
|---|---|---|
| Shift by \(z_{\max} = 2.0\) | \(z - z_{\max}\) | \([-1.0, \; 0.0, \; -1.5]\) |
| Exponentials | \(e^{z - z_{\max}}\) | \([0.3679, \; 1.0, \; 0.2231]\) |
| Normalise (sum \(1.5910\)) | \(a = e^{z - z_{\max}} / 1.5910\) | \([0.2312, \; 0.6285, \; 0.1402]\) |
| Dot product | \(\langle g, a \rangle\) | \(-0.0423\) |
| Backward | \(\delta_i = a_i (g_i - \langle g, a \rangle)\) | \([0.0792, \; -0.0991, \; 0.0200]\) |
The outputs sum to \(1\) and the gradient sums to \(0\) (\(\sum_i \delta_i = \langle g, a \rangle - \langle g, a \rangle \sum_i a_i = 0\)), reflecting the shift invariance.
| Variant | Key Difference |
|---|---|
| Identity (linear) | \(f(z) = z\), \(f' = 1\); raw regression outputs and logits |
| PReLU (Parametric ReLU) | \(\alpha\) is a learnable parameter per channel |
| ELU (Exponential LU) | \(\alpha(e^z - 1)\) for \(z < 0\); smooth and zero-centered |
| GELU (Gaussian Error LU) | \(z \cdot \Phi(z)\); used in Transformers |
| Swish / SiLU | \(z \cdot \sigma(z)\); smooth, non-monotonic |
| Hard Sigmoid / Hard Tanh | Piece-wise linear approximations; no transcendentals |
graph TD
Act["Activation Functions"]
Layer["Dense Layer"]
NN["Neural Network"]
Loss["Loss Functions"]
Act --> Layer
Layer --> NN
NN --> Loss
| Component | Relationship |
|---|---|
| Dense Layer | Applies the activation after the affine transformation and calls its backward pass |
| Neural Network | Activations enable the non-linear function approximation that makes deep networks useful |
| Loss Functions | Sigmoid output feeds binary cross-entropy; categorical cross-entropy takes logits and applies softmax itself — no Softmax layer |
A dense (fully-connected) layer is the most fundamental building block of a neural network. It maps an input vector \(a_{\text{in}} \in \mathbb{R}^n\) to an output vector \(a_{\text{out}} \in \mathbb{R}^m\) through a learnable affine transformation followed by an activation function:
\[a_{\text{out}} = f(W \, a_{\text{in}} + b)\]
Every input neuron is connected to every output neuron — hence “fully connected.” The layer’s parameters are the weight matrix \(W\) and bias vector \(b\); training adjusts these to minimize the loss.
In this library, input size, output size, and parameter count are all compile-time constants, enabling static allocation and dimension checking with zero runtime overhead. The caller supplies the initial weights; the biases start at zero.
\[z = W \, a_{\text{in}} + b, \qquad a_{\text{out}} = f(z)\]
where \(W \in \mathbb{R}^{m \times n}\), \(b \in \mathbb{R}^m\), and \(f\) is the activation function. The layer caches \(a_{\text{in}}\), \(z\) and \(a_{\text{out}}\) for the backward pass.
Given the upstream gradient \(g = \partial \mathcal{L} / \partial a_{\text{out}} \in \mathbb{R}^m\) for one sample:
Pre-activation gradient — the transposed activation Jacobian applied to \(g\): \[\delta = \frac{\partial \mathcal{L}}{\partial z} = J_f(z)^T g\] For an element-wise activation the Jacobian is diagonal and this reduces to \(\delta = g \odot f'(z)\). For Softmax it is \(\delta_i = a_i \bigl(g_i - \sum_k g_k a_k\bigr)\) with \(a = a_{\text{out}}\), still \(O(m)\) (see Softmax).
Weight gradient: \[\frac{\partial \mathcal{L}}{\partial W} = \delta \, a_{\text{in}}^T\]
Bias gradient: \[\frac{\partial \mathcal{L}}{\partial b} = \delta\]
Input gradient (propagated to the previous layer): \[\frac{\partial \mathcal{L}}{\partial a_{\text{in}}} = W^T \delta\]
The backward pass requires a preceding forward pass (it uses the cached \(a_{\text{in}}\), \(z\), \(a_{\text{out}}\)). Each call overwrites the stored parameter gradients with those of the current sample; they are not accumulated and are not yet exposed to an optimizer (roadmap N8).
The parameters are stored as one flat vector: \(W\) in row-major order (row \(i\) holds the weights of output neuron \(i\)) followed by \(b\):
\[\theta = [\, W_{1,1}, \ldots, W_{1,n}, \; W_{2,1}, \ldots, W_{m,n}, \; b_1, \ldots, b_m \,], \qquad P = m\,n + m = m(n + 1)\]
For a layer with 128 inputs and 64 outputs: \(P = 64 \times 129 = 8{,}256\) parameters.
| Operation | Time | Extra space |
|---|---|---|
| Forward (\(W a + b\), then \(f\)) | \(O(m \cdot n)\) | caches \(a_{\text{in}}\) (\(n\)), \(z\) (\(m\)), \(a_{\text{out}}\) (\(m\)) |
| Backward (\(\delta\), \(\nabla W\), \(\nabla b\), \(W^T\delta\)) | \(O(m \cdot n)\) | \(O(m)\) temporary \(\delta\) |
The matrix-vector products dominate both passes. For embedded networks (e.g. \(n = 32, m = 16\)), a single forward pass takes 512 multiply-accumulate operations.
Resident memory per layer. The object holds the parameters (\(P\) values), a parameter-gradient buffer of the same size (\(P\)), and the cached input, pre-activation, output and input-gradient vectors (\(2n + 2m\)):
\[\text{values} = 2P + 2n + 2m = 2mn + 4m + 2n\]
That is roughly twice the weight matrix. A \(256 \times 256\) float layer holds \(2 \times 65{,}792 + 1{,}024 = 132{,}608\) values, about \(518\,\text{KiB}\) — not the \(256\,\text{KiB}\) of the weights alone. The model adds a transient \(P\)-sized copy when parameters are exported or imported.
Layer: 3 inputs → 2 outputs, ReLU activation.
\[W = \begin{bmatrix} 0.5 & -0.3 & 0.8 \\ 0.1 & 0.7 & -0.2 \end{bmatrix}, \quad b = \begin{bmatrix} 0.1 \\ -0.1 \end{bmatrix}, \quad a_{\text{in}} = \begin{bmatrix} 1.0 \\ 0.5 \\ -1.0 \end{bmatrix}\]
Forward:
| Step | Computation | Result |
|---|---|---|
| \(z = W a_{\text{in}} + b\) | \([0.5 - 0.15 - 0.8 + 0.1,\; 0.1 + 0.35 + 0.2 - 0.1]\) | \([-0.35,\; 0.55]^T\) |
| \(a_{\text{out}} = \text{ReLU}(z)\) | \([\max(0, -0.35),\; \max(0, 0.55)]\) | \([0.0,\; 0.55]^T\) |
Backward with \(g = \partial \mathcal{L} / \partial a_{\text{out}} = [0.2,\; -0.4]^T\):
| Step | Computation | Result |
|---|---|---|
| \(\delta = g \odot \text{ReLU}'(z)\) | \([0.2 \cdot 0,\; -0.4 \cdot 1]\) | \([0,\; -0.4]^T\) |
| \(\nabla W = \delta \, a_{\text{in}}^T\) | row 1: all zeros; row 2: \(-0.4 \times [1, 0.5, -1]\) | \(\begin{bmatrix}0 & 0 & 0\\-0.4 & -0.2 & 0.4\end{bmatrix}\) |
| \(\nabla b = \delta\) | — | \([0,\; -0.4]^T\) |
| \(\nabla a_{\text{in}} = W^T \delta\) | \(-0.4 \times [0.1,\; 0.7,\; -0.2]\) | \([-0.04,\; -0.28,\; 0.08]^T\) |
In the flat parameter layout, \(\nabla\theta = [0, 0, 0, -0.4, -0.2, 0.4, 0, -0.4]\).
| Variant | Key Difference |
|---|---|
| Convolutional layer | Weight sharing across spatial positions; \(O(k^2 \cdot c)\) parameters per filter instead of \(O(n \cdot m)\) |
| Recurrent layer | Shares weights across time steps; adds a hidden state feedback connection |
| Batch normalization layer | Normalizes activations to zero mean and unit variance; accelerates training |
| Dropout layer | Randomly zeros activations during training; regularization effect |
| Sparse layer | Only a subset of connections exist; reduces parameter count and computation |
graph TD
Layer["Dense Layer"]
Act["Activation Functions"]
Model["Model"]
Opt["Optimizer"]
LR["Linear Regression"]
Act --> Layer
Layer --> Model
Model --> Opt
Layer -.->|"identity activation, MSE loss"| LR
| Component | Relationship |
|---|---|
| Activation Functions | Applied after the affine transformation; its Jacobian-vector product gives \(\delta\) |
| Model | Chains multiple dense layers and concatenates their flat parameter vectors |
| Optimizer (numerical-toolbox-cpp) | Updates the model’s flat parameter vector; feeding it the layer’s own gradients is roadmap N8/N16 |
| Linear Regression (numerical-toolbox-cpp) | A dense layer with an identity activation \(f(z) = z\) and MSE loss is linear regression; the identity activation is roadmap N1 |
A loss function \(\mathcal{L}(\hat{y}, y)\) quantifies how far a model’s prediction \(\hat{y}\) is from the true target \(y\). Training a neural network means finding parameters \(\theta\) that minimize the expected loss over the training data:
\[\theta^* = \arg\min_\theta \; \mathbb{E}[\mathcal{L}(f_\theta(x), y)]\]
The loss function defines the entire learning objective — different losses lead to different optimal models even on the same data. It must also provide the gradient \(\nabla_{\hat{y}} \mathcal{L}\) that seeds back-propagation.
In this library every loss is an objective over one vector argument \(v \in \mathbb{R}^N\) — the \(N\) outputs of one sample — compared against a target \(y \in \mathbb{R}^N\) fixed when the loss is created. A regularization term \(R\) from numerical-toolbox-cpp is added to the same argument:
\[J(v) = \mathcal{L}(v, y) + R(v), \qquad \nabla J(v) = \nabla_v \mathcal{L}(v, y) + \nabla R(v)\]
For MSE, MAE and BCE the argument is the prediction \(v = \hat{y}\); for categorical cross-entropy it is the vector of logits \(v = z\).
All data terms are means over the \(N\) outputs, so their gradients carry a \(1/N\) factor.
\[\mathcal{L}_{\text{MSE}} = \frac{1}{N} \sum_{i=1}^{N} (\hat{y}_i - y_i)^2, \qquad \frac{\partial \mathcal{L}}{\partial \hat{y}_i} = \frac{2}{N}(\hat{y}_i - y_i)\]
\[\mathcal{L}_{\text{MAE}} = \frac{1}{N} \sum_{i=1}^{N} |\hat{y}_i - y_i|, \qquad \frac{\partial \mathcal{L}}{\partial \hat{y}_i} = \frac{1}{N} \operatorname{sign}(\hat{y}_i - y_i)\]
The input is a vector of probabilities \(p = \hat{y} \in [0, 1]^N\) (e.g. a Sigmoid output) and \(y_i \in [0, 1]\). Each probability is first clamped to \(\tilde{p}_i = \operatorname{clamp}(p_i, \varepsilon, 1 - \varepsilon)\) with \(\varepsilon = 10^{-7}\):
\[\mathcal{L}_{\text{BCE}} = -\frac{1}{N}\sum_{i=1}^{N} \left[ y_i \log \tilde{p}_i + (1 - y_i) \log(1 - \tilde{p}_i) \right]\]
\[\frac{\partial \mathcal{L}}{\partial p_i} = \frac{1}{N} \cdot \frac{\tilde{p}_i - y_i}{\tilde{p}_i(1 - \tilde{p}_i)}\]
The input is a vector of logits \(z \in \mathbb{R}^k\) (\(N = k\) classes), not probabilities. The loss fuses softmax and cross-entropy:
\[\mathcal{L}_{\text{CCE}} = -\sum_{i=1}^{k} y_i \log \operatorname{softmax}(z)_i = \Bigl(\sum_{i} y_i\Bigr) \operatorname{LSE}(z) - \sum_{i} y_i z_i, \qquad \operatorname{LSE}(z) = \log \sum_{j} e^{z_j}\]
\[\frac{\partial \mathcal{L}}{\partial z_i} = \operatorname{softmax}(z)_i \sum_{j} y_j - y_i\]
| Loss | Cost | Gradient | Transcendentals |
|---|---|---|---|
| MSE | \(O(N)\) | \(O(N)\) | none |
| MAE | \(O(N)\) | \(O(N)\) | none (sign) |
| BCE | \(O(N)\) | \(O(N)\) | \(2N\) logarithms (cost only) |
| CCE | \(O(k)\) | \(O(k)\) | \(k\) exponentials + 1 logarithm per call |
All losses are linear in the output dimension and use \(O(1)\) extra space besides the returned gradient; the regularization term adds its own \(O(N)\). The cost is negligible compared to the dense-layer matrix products.
Categorical cross-entropy on logits. 3 classes, target \(y = [0, 1, 0]\) (class 2), logits \(z = [1.0, 2.0, 0.5]\).
| Step | Computation | Result |
|---|---|---|
| Max shift | \(z_{\max} = 2.0\), \(z - z_{\max}\) | \([-1.0,\; 0.0,\; -1.5]\) |
| Shifted exponentials | \(e^{z - z_{\max}}\) | \([0.3679,\; 1.0,\; 0.2231]\), sum \(1.5910\) |
| Log-sum-exp | \(2.0 + \log 1.5910\) | \(2.4644\) |
| Cost | \(\operatorname{LSE}(z) \cdot 1 - z_2\) | \(2.4644 - 2.0 = 0.4644\) |
| Softmax | \(e^{z_i - \operatorname{LSE}(z)}\) | \([0.2312,\; 0.6285,\; 0.1402]\) |
| Gradient | \(\operatorname{softmax}(z) - y\) | \([0.2312,\; -0.3715,\; 0.1402]\) |
The cost equals \(-\log 0.6285\), the negative log-probability of the true class. The gradient pushes the true logit up and the others down, and sums to zero.
For comparison — MSE on the same probabilities \(\hat{y} = [0.2312, 0.6285, 0.1402]\):
\[\mathcal{L}_{\text{MSE}} = \frac{1}{3}\bigl[0.2312^2 + 0.3715^2 + 0.1402^2\bigr] = 0.0704, \qquad \nabla = \frac{2}{3}(\hat{y} - y) = [0.1541,\; -0.2476,\; 0.0935]\]
Binary cross-entropy with \(p = [0.9, 0.2]\) and \(y = [1, 0]\) (\(N = 2\), no clamping active):
\[\mathcal{L}_{\text{BCE}} = -\tfrac{1}{2}\bigl[\log 0.9 + \log 0.8\bigr] = 0.1643, \qquad \nabla = \tfrac{1}{2}\Bigl[\tfrac{0.9 - 1}{0.9 \cdot 0.1},\; \tfrac{0.2 - 0}{0.2 \cdot 0.8}\Bigr] = [-0.5556,\; 0.625]\]
| Variant | Key Difference |
|---|---|
| Huber loss | Quadratic for small errors, linear for large; robust regression |
| Logit-input BCE | Fuses Sigmoid and BCE: \(\operatorname{softplus}(z) - yz\); no clamp needed |
| Focal loss | Down-weights well-classified examples; addresses class imbalance |
| KL divergence | Measures distance between two distributions; used in variational inference |
| Hinge loss | Margin-based; used in SVMs and some neural classifiers |
| Contrastive loss | Learns similarity metrics; used in Siamese networks |
graph TD
Loss["Loss Functions"]
Act["Activation Functions"]
Opt["Optimizer"]
Reg["Regularization"]
Model["Model"]
LR["Linear Regression"]
Act -->|"output activation must match loss"| Loss
Loss --> Opt
Reg -->|"added to loss"| Loss
Loss --> Model
Loss -.->|"MSE + normal equation"| LR
| Component | Relationship |
|---|---|
| Activation Functions | Output activation must match: Sigmoid ↔︎ BCE, identity (logits) ↔︎ CCE, identity ↔︎ MSE/MAE |
| Model | The model’s training step minimises an objective of this form over the flat parameter vector \(\theta\) |
| Optimizer (numerical-toolbox-cpp) | Minimises \(J\) using its cost and gradient |
| Regularization (numerical-toolbox-cpp) | Adds \(R(v)\) and \(\nabla R(v)\) to the data term |
| Linear Regression (numerical-toolbox-cpp) | Solved analytically when the loss is MSE and the model is linear |
A neural network is a parameterized function \(f_\theta: \mathbb{R}^n \to \mathbb{R}^m\) built by composing simple, differentiable transformations called layers. Each layer applies an affine map followed by a non-linear activation, and the whole composition is trained by gradient-based optimization.
Neural networks are powerful because of the universal approximation theorem: a single hidden layer with enough neurons can approximate any continuous function on a compact set to arbitrary accuracy. In practice, depth (many layers) is more parameter-efficient than width for learning hierarchical features.
This library provides a minimal, statically-sized neural network framework designed for embedded inference — no heap allocation, no dynamic shapes, full compile-time dimension checking. The forward and backward passes described below are implemented; the parameter update loop that turns them into on-device training is roadmap N8/N9/N16 (see Model).
Given \(L\) layers, the network computes:
\[a_0 = x\] \[z_\ell = W_\ell \, a_{\ell-1} + b_\ell, \quad \ell = 1, \ldots, L\] \[a_\ell = f_\ell(z_\ell)\] \[\hat{y} = a_L\]
where \(W_\ell \in \mathbb{R}^{n_\ell \times n_{\ell-1}}\) are weights, \(b_\ell \in \mathbb{R}^{n_\ell}\) are biases, and \(f_\ell\) is the activation function for layer \(\ell\).
Training minimizes a scalar loss \(\mathcal{L}(\hat{y}, y)\) that measures how far the prediction \(\hat{y}\) is from the target \(y\). Common choices (see Loss Functions):
| Loss | Formula | Input | Use Case |
|---|---|---|---|
| MSE | \(\frac{1}{N}\sum_i(\hat{y}_i - y_i)^2\) | predictions | Regression |
| BCE | \(-\frac{1}{N}\sum_i[y_i \log p_i + (1-y_i)\log(1-p_i)]\) | probabilities \(p\) | Binary classification |
| CCE | \(-\sum_i y_i \log \operatorname{softmax}(z)_i\) | logits \(z\) | Multi-class classification |
Backpropagation computes \(\nabla_\theta \mathcal{L}\) via the chain rule, working from the output layer backward. With \(g_L = \nabla_{a_L} \mathcal{L}\) and \(J_{f_\ell}\) the Jacobian of the activation:
\[\delta_\ell = J_{f_\ell}(z_\ell)^T g_\ell, \qquad g_{\ell-1} = W_\ell^T \delta_\ell\]
For element-wise activations \(J_{f_\ell}\) is diagonal and \(\delta_\ell = g_\ell \odot f_\ell'(z_\ell)\); for Softmax it is the \(O(n_\ell)\) product \(\delta_\ell = a_\ell \odot (g_\ell - \langle g_\ell, a_\ell \rangle)\) (see Activation Functions). The gradients with respect to the parameters are:
\[\frac{\partial \mathcal{L}}{\partial W_\ell} = \delta_\ell \, a_{\ell-1}^T, \qquad \frac{\partial \mathcal{L}}{\partial b_\ell} = \delta_\ell\]
An optimizer uses the gradients to update the parameter vector \(\theta\):
\[\theta_{t+1} = \theta_t - \eta \, \nabla_\theta \mathcal{L}\]
where \(\eta\) is the learning rate. More sophisticated optimizers (momentum, Adam) modify this basic rule. With a mini-batch of \(B\) samples, \(\nabla_\theta \mathcal{L}\) is the average of the per-sample gradients.
| Phase | Time | Space |
|---|---|---|
| Forward pass | \(O\!\left(\sum_{\ell=1}^L n_\ell \cdot n_{\ell-1}\right)\) | \(O\!\left(\sum_\ell n_\ell\right)\) activations |
| Backward pass | Same as forward | Same + \(O(P)\) gradient storage |
| Parameter update | \(O(P)\) | \(O(P)\) optimizer state |
where \(P = \sum_\ell (n_\ell \cdot n_{\ell-1} + n_\ell)\) is the total parameter count. Both passes cost about one multiply-accumulate per weight, so a network with \(P = 10{,}000\) parameters needs on the order of \(10^4\) MACs per sample and pass.
Network: 2 inputs → 2 hidden (ReLU) → 1 output (Sigmoid), trained with binary cross-entropy on one XOR sample.
Architecture:
graph LR
x1((x₁)) --> h1((h₁))
x1 --> h2((h₂))
x2((x₂)) --> h1
x2 --> h2
h1 --> y((ŷ))
h2 --> y
Initial parameters:
\[W_1 = \begin{bmatrix} 0.2 & 0.4 \\ -0.3 & 0.1 \end{bmatrix}, \quad b_1 = \begin{bmatrix} 0.1 \\ 0.2 \end{bmatrix}, \quad W_2 = \begin{bmatrix} 0.5 & -0.4 \end{bmatrix}, \quad b_2 = [0]\]
Forward pass with input \(x = [1, 0]^T\), target \(y = 1\):
| Step | Computation | Result |
|---|---|---|
| Hidden pre-activation | \(z_1 = W_1 x + b_1\) | \([0.3, -0.1]^T\) |
| Hidden activation | \(a_1 = \text{ReLU}(z_1)\) | \([0.3, 0.0]^T\) |
| Output pre-activation | \(z_2 = W_2 a_1 + b_2\) | \(0.15\) |
| Output activation | \(\hat{y} = \sigma(z_2)\) | \(0.5374\) |
| Loss | \(\mathcal{L} = -(y\log\hat{y} + (1-y)\log(1-\hat{y}))\) | \(0.6210\) |
Backward pass (Sigmoid followed by BCE simplifies to \(\delta_2 = \hat{y} - y\)):
| Step | Computation | Result |
|---|---|---|
| Output gradient | \(\delta_2 = \hat{y} - y\) | \(-0.4626\) |
| \(\nabla W_2\) | \(\delta_2 \, a_1^T\) | \([-0.1388, 0]\) |
| \(\nabla b_2\) | \(\delta_2\) | \(-0.4626\) |
| Hidden gradient | \(\delta_1 = (W_2^T \delta_2) \odot \text{ReLU}'(z_1)\) | \([-0.2313, 0]^T\) |
| \(\nabla W_1\) | \(\delta_1 \, x^T\) | \([[-0.2313, 0], [0, 0]]\) |
| \(\nabla b_1\) | \(\delta_1\) | \([-0.2313, 0]^T\) |
Update with \(\eta = 0.1\): \(W_2 \leftarrow [0.5139, -0.4]\), \(b_2 \leftarrow 0.0463\), \(W_{1,11} \leftarrow 0.2231\), \(b_{1,1} \leftarrow 0.1231\) (all other parameters have zero gradient). Repeating the forward pass gives \(\hat{y} = \sigma(0.2242) = 0.5558\) and \(\mathcal{L} = 0.5873\) — the loss decreased. Training repeats this step over all four XOR samples.
| Variant | Key Difference |
|---|---|
| Convolutional Neural Network (CNN) | Layers share weights spatially; efficient for image/signal data |
| Recurrent Neural Network (RNN) | Layers share weights across time steps; models sequences |
| Residual Network (ResNet) | Skip connections mitigate vanishing gradients in very deep networks |
| Transformer | Attention-based; no recurrence; state-of-the-art for sequences |
| Quantized Neural Network | Weights and activations in low-bit integers; smaller and faster on MCUs |
graph TD
NN["Neural Network"]
Layer["Dense Layer"]
Act["Activation Functions"]
Loss["Loss Functions"]
Opt["Optimizer"]
Reg["Regularization"]
Model["Model"]
LR["Linear Regression"]
Layer --> NN
Act --> NN
Loss --> NN
Opt --> NN
Reg --> NN
NN --> Model
NN -.->|"single layer, identity activation, MSE loss"| LR
| Component | Relationship |
|---|---|
| Dense Layer | The fundamental building block; computes affine transformations |
| Activation Functions | Introduce non-linearity after each layer |
| Loss Functions | Define the training objective |
| Optimizer (numerical-toolbox-cpp) | Drives parameter updates via gradient descent |
| Regularization (numerical-toolbox-cpp) | Penalizes complexity to prevent overfitting |
| Model | Composes layers into one network with a flat parameter vector |
| Linear Regression (numerical-toolbox-cpp) | Special case: single layer, identity activation, MSE loss |
A model composes a sequence of layers — today dense layers — into a single function \(f_\theta: \mathbb{R}^n \to \mathbb{R}^m\). It:
The model is fully statically typed — layer dimensions, parameter counts, and memory footprints are all known at compile time, enabling zero-overhead abstraction on embedded targets.
For \(L\) layers with transformations \(f_1, f_2, \ldots, f_L\):
\[\hat{y} = (f_L \circ f_{L-1} \circ \cdots \circ f_1)(x) = f_L(f_{L-1}(\ldots f_1(x) \ldots))\]
Each \(f_\ell\) is a dense layer: \(a_\ell = f_\ell(a_{\ell-1}) = \sigma_\ell(W_\ell a_{\ell-1} + b_\ell)\) with \(a_0 = x\), \(W_\ell \in \mathbb{R}^{m_\ell \times n_\ell}\) and \(n_{\ell+1} = m_\ell\).
All weights and biases are concatenated in layer order; each \(W_\ell\) is flattened row-major (row \(i\) = weights of output neuron \(i\)), followed by its bias:
\[\theta = [\, \operatorname{rows}(W_1),\ b_1,\ \operatorname{rows}(W_2),\ b_2,\ \ldots,\ \operatorname{rows}(W_L),\ b_L \,] \in \mathbb{R}^P, \qquad P = \sum_{\ell=1}^L m_\ell(n_\ell + 1)\]
Export and import use exactly this order, so exporting and re-importing is the identity.
graph LR
X["x ∈ ℝⁿ"] --> L1["Layer 1"] --> L2["Layer 2"] --> Ldots["⋯"] --> LL["Layer L"] --> Y["ŷ ∈ ℝᵐ"]
Given \(g_L = \partial \mathcal{L} / \partial \hat{y}\), the layers are visited in reverse. Layer \(\ell\) receives \(g_\ell = \partial \mathcal{L} / \partial a_\ell\) and computes
\[\delta_\ell = J_{\sigma_\ell}(z_\ell)^T g_\ell, \qquad \nabla W_\ell = \delta_\ell\, a_{\ell-1}^T, \qquad \nabla b_\ell = \delta_\ell, \qquad g_{\ell-1} = W_\ell^T \delta_\ell\]
so each layer applies its own activation derivative before handing \(g_{\ell-1}\) to its predecessor. The model returns \(g_0 = \partial \mathcal{L} / \partial x\). The per-layer parameter gradients stay inside the layers; exporting them as a flat \(\nabla_\theta \mathcal{L}\) is roadmap N8.
The training step takes an optimizer, an objective \(J: \mathbb{R}^P \to \mathbb{R}\) with its gradient \(\nabla J\), and a starting point \(\theta_0\):
\[\theta^* = \operatorname{Optimize}(J, \theta_0) \approx \arg\min_\theta J(\theta), \qquad \text{then load } \theta^* \text{ into the layers}\]
It does not iterate over a dataset and does not call the forward or backward pass: the objective must itself encode whatever it measures about \(\theta\) (for example a user-written empirical risk). If a per-sample loss is passed as \(J\), its argument is \(\theta\) and its target is a parameter vector — it regresses the parameters onto that target. End-to-end back-propagation training — per-sample loss on \(\hat{y}\) (N9), exported parameter gradients (N8) and a mini-batch gradient step (N16) — is on the roadmap.
graph LR
J["Objective J(θ), ∇J(θ)"] --> OPT["Optimizer from θ₀"] --> TH["θ*"] --> SET["Load θ* into layers"]
| Operation | Time | Space |
|---|---|---|
| Forward pass | \(O(P)\) | per-layer caches \(O(\sum_\ell (n_\ell + m_\ell))\) |
| Backward pass | \(O(P)\) | per-layer gradient buffers \(O(P)\) (resident) |
| Parameter export | \(O(P)\) | \(O(P)\) flat vector returned |
| Parameter import | \(O(P)\) | \(O(\max_\ell P_\ell)\) transient per-layer copy |
| Training step | optimizer-dependent | optimizer state + \(O(P)\) |
Forward and backward are linear in the total parameter count \(P\) (one multiply-accumulate per weight). The resident model is the sum of its layers — about \(2P\) values plus the activation caches for dense layers (see Dense Layer).
Model: 2 → 3 → 1. Layer 1 uses ReLU, layer 2 uses Sigmoid.
Compile-time verification chain:
| Check | Condition | Status |
|---|---|---|
| At least one layer | \(L = 2 \ge 1\) | ✓ |
| Layer 1 input size = model input size | \(2 = 2\) | ✓ |
| Layer 1 output size = layer 2 input size | \(3 = 3\) | ✓ |
| Layer 2 output size = model output size | \(1 = 1\) | ✓ |
| Every layer shares the model’s value type | float | ✓ |
Parameters (\(P = 3(2 + 1) + 1(3 + 1) = 13\)):
\[W_1 = \begin{bmatrix} -0.5 & 0.2 \\ 0.4 & 0.6 \\ 0.1 & 0.2 \end{bmatrix}, \quad b_1 = \begin{bmatrix} 0.1 \\ 0.1 \\ 0.1 \end{bmatrix}, \quad W_2 = \begin{bmatrix} 0.3 & 0.5 & 0.2 \end{bmatrix}, \quad b_2 = [0.03]\]
| Index | Parameter | Values |
|---|---|---|
| 0–5 | \(W_1\) row-major (\(3 \times 2\)) | \(-0.5, 0.2, 0.4, 0.6, 0.1, 0.2\) |
| 6–8 | \(b_1\) | \(0.1, 0.1, 0.1\) |
| 9–11 | \(W_2\) row-major (\(1 \times 3\)) | \(0.3, 0.5, 0.2\) |
| 12 | \(b_2\) | \(0.03\) |
Forward pass with \(x = [1.0, 0.5]^T\):
| Step | Computation | Result |
|---|---|---|
| Layer 1 | \(z_1 = W_1 x + b_1\) | \([-0.3,\; 0.8,\; 0.3]^T\) |
| \(a_1 = \text{ReLU}(z_1)\) | \([0.0,\; 0.8,\; 0.3]^T\) | |
| Layer 2 | \(z_2 = W_2 a_1 + b_2 = 0.4 + 0.06 + 0.03\) | \(0.49\) |
| \(\hat{y} = \sigma(z_2)\) | \(0.6201\) |
Backward pass with loss gradient \(g_2 = \partial \mathcal{L} / \partial \hat{y} = [0.12]\):
| Step | Computation | Result |
|---|---|---|
| Layer 2 | \(\delta_2 = g_2 \, \sigma'(z_2) = 0.12 \cdot 0.6201 \cdot 0.3799\) | \(0.02827\) |
| \(\nabla W_2 = \delta_2 a_1^T\), \(\nabla b_2 = \delta_2\) | \([0,\; 0.02262,\; 0.00848]\), \(0.02827\) | |
| \(g_1 = W_2^T \delta_2\) (handed to layer 1) | \([0.00848,\; 0.01413,\; 0.00565]^T\) | |
| Layer 1 | \(\delta_1 = g_1 \odot \text{ReLU}'(z_1)\) | \([0,\; 0.01413,\; 0.00565]^T\) |
| \(\nabla W_1 = \delta_1 x^T\), \(\nabla b_1 = \delta_1\) | rows \([0, 0]\), \([0.01413, 0.00707]\), \([0.00565, 0.00283]\) | |
| \(g_0 = W_1^T \delta_1\) (returned by the model) | \([0.00622,\; 0.00961]^T\) |
Note where \(\text{ReLU}'(z_1)\) enters: inside layer 1’s own backward pass, not in layer 2’s.
| Variant | Key Difference |
|---|---|
| Sequential model (dynamic) | Layers stored in a container; dimension checked at runtime instead of compile time |
| Functional API | Supports branching and merging (DAG topology instead of linear chain) |
| Residual model | Adds skip connections: \(a_{\ell+2} = f_{\ell+1}(a_\ell) + a_\ell\) |
| Recurrent model | Unrolls the same layer across time steps |
graph TD
Model["Model"]
Layer["Dense Layer"]
Loss["Loss Functions"]
Opt["Optimizer"]
Reg["Regularization"]
NN["Neural Network"]
Layer --> Model
Model --> Opt
Loss --> Opt
Reg --> Loss
Model --> NN
| Component | Relationship |
|---|---|
| Dense Layer | The model is a fixed sequence of layers evaluated in order |
| Neural Network | Overview of forward propagation, back-propagation and the parameter update the model builds on |
| Loss Functions | Objectives of the same form; the training step accepts one sized to the parameter vector |
| Optimizer (numerical-toolbox-cpp) | Minimises the objective over the flat parameter vector and returns the minimiser |
| Regularization (numerical-toolbox-cpp) | Added to the objective’s argument by the loss |