In Brief
For executives and procurement. This article documents how a small neural network is trained and run without a machine learning library. The same reasoning behind Wholesale Dito Store's zero-dependency stack also applies to AI, and the article explains why AI is excluded from the customer data path and the calculation path. That policy is documented in Why We Built a Zero-Dependency Stack.
For developers. The article walks through a complete feedforward neural network, ten inputs, six hidden neurons, one output, trained with hand-written backpropagation in PHP and run in the browser in JavaScript. Full code for the training loop and the inference function is included.
What this article is not. It is not a production spam filter, and it is not a case for putting AI in your procurement workflow. It is a documented demonstration of the underlying mechanics.
Direct Answer
AI works by adjusting numbers. A neural network is a stack of matrix multiplications with a non-linear function between the layers. Training is the process of nudging those numbers until the output matches the label. Inference is the process of running the trained numbers on new input to produce a prediction.
The article below trains a narrow AI from scratch, using only PHP for training and vanilla JavaScript for inference, and shows what happens at every step. No machine learning library. No Python. No cloud service. Just the arithmetic that has been under every neural network since the 1950s, applied to a single, well-defined task.
Who This Is For
This article is written for three audiences at once:
- Developers and engineers who want to see a neural network built without a machine learning framework, and who want to understand what backpropagation is actually doing.
- Business owners who want to understand what AI is and is not, before deciding whether to buy it, build it, or avoid it.
- Procurement teams and executives who need to know how a technology supplier thinks about AI, and what that means for data handling and vendor risk.
The article is split into two parts. Part 1 is the technical walkthrough. Part 2 is the strategic decision. Both are visible. Neither is hidden behind the other.
Why This Matters
The question "how does AI actually work" is often answered with metaphors. A neural network is described as a brain, a model is described as learning, and the whole thing is wrapped in language that makes it sound alive. It is not alive. It is arithmetic. That is the entire point.
When the arithmetic is visible, three things become clear. First, AI is not magic. Second, AI is a dependency. Third, AI's limits are structural, not accidental. Those three things are the reason this article documents the mechanic, and the reason the company keeps AI out of the customer data path.
At Wholesale Dito Store, our technical co-founder first encountered the raw material of AI in 2010, through the company's work-from-home data entry operation, in the same broader economy where ordinary people were generating the labeled data that machine learning would eventually consume. In 2012, operations transitioned to full-scale web and software development. In 2016, the same person began building narrow AI systems of his own. The interest faded before 2020, and the work stopped. This article is a one-time return to that form, to document what the mechanic actually looks like.
Part 1: The Technical Walkthrough
For developers and engineers. If you are here for the strategic decision behind our AI policy, skip to Part 2 below.
The next sections explain what a narrow AI is, what a neural network actually contains, and how training and inference work step by step, with real code.
What "Narrow AI" Means
A narrow AI is a system that performs one task. It does not generalize. It does not have opinions. It does not know anything outside the data it was trained on.
The spam filter described in this article classifies text as spam or not spam. It has no understanding of the words it reads. It has no memory of previous messages. It cannot tell you what spam is. It has one ability: given ten input numbers, produce one output number.
That is the whole thing. When people say "AI," they usually mean a system like this, scaled up by many orders of magnitude and trained on vastly more data. The mechanics do not change. The scale does.
What a Neural Network Actually Is
Strip away the marketing and a neural network is three things:
- A set of numbers called weights.
- A set of numbers called biases.
- A formula that combines inputs, weights, and biases, and squashes the result into a range.
For a simple feedforward network, the formula is matrix multiplication, then a non-linear function, then more matrix multiplication, then another non-linear function. Repeat for as many layers as you have.
The non-linear function is what makes the network interesting. Without it, stacking layers would just produce another linear function, and the whole thing would collapse into a single layer. Common choices are sigmoid, tanh, and ReLU.
Why PHP and JavaScript
PHP and JavaScript are not the usual tools for machine learning. Python is, because Python has NumPy, PyTorch, and scikit-learn, which handle the arithmetic at a speed PHP cannot match.
But the arithmetic itself is not hard. A network with ten inputs, six hidden neurons, and one output neuron has seventy-three numbers to adjust. PHP can do that in a fraction of a second. The reason to use PHP and JavaScript is not performance. It is to show the mechanics without the abstraction of a library.
If you understand the PHP version, you understand what PyTorch is doing. The library just does it faster and on a much larger scale.
The Task: Classify a Message as Spam or Not Spam
We pick a narrow problem. A message is either spam or not spam. The vocabulary is ten words: free, win, money, click, offer, meeting, project, report, hello, invoice.
The training data is six labeled examples. Three are spam, three are not. Real spam filters train on millions of messages. Six is enough to demonstrate the mechanics.
Each message becomes a vector of ten numbers. A 1 means the word is present. A 0 means it is not. For example, the message "free money click offer" becomes [1, 0, 1, 1, 1, 0, 0, 0, 0, 0].
The network receives that vector and outputs a single number between 0 and 1. Above 0.5 means spam. Below 0.5 means not spam.
Step 1: The Architecture
The network has three layers:
- Input layer: ten nodes, one for each vocabulary word.
- Hidden layer: six neurons, chosen arbitrarily.
- Output layer: one neuron, producing the spam probability.
The weights connecting the input layer to the hidden layer form a 6 by 10 matrix called W1. The biases for the hidden neurons are a vector of six numbers called b1.
The weights connecting the hidden layer to the output neuron form a 1 by 6 matrix called W2. The output bias is a single number called b2.
At the start of training, these numbers are random. Training will adjust them.
Step 2: Forward Propagation
Forward propagation is how the network produces a prediction. It happens in two stages.
Stage one, the hidden layer. For each of the six hidden neurons:
- Multiply each input by its corresponding weight.
- Add the bias.
- Apply the sigmoid function.
Sigmoid squashes any number into a range between 0 and 1. A neuron that sums to a large positive number outputs a value close to 1. A neuron that sums to a large negative number outputs a value close to 0. A neuron whose sum is near zero outputs something close to 0.5.
Stage two, the output layer. Take the six hidden activations, multiply each by its output weight, add the output bias, and apply sigmoid again. The result is the final prediction.
Step 3: Measuring the Error
Compare the prediction to the actual label. If the message is spam, the label is 1. If not, the label is 0. The difference between the prediction and the label is the error.
The goal of training is to reduce that error across all training examples.
Step 4: Backpropagation
Backpropagation is how the network learns. It is an application of the chain rule from calculus, but the concept is simpler than the name suggests.
The network calculates how much each weight and bias contributed to the error. Then it adjusts each one slightly in the direction that reduces the error. The size of the adjustment is controlled by a number called the learning rate.
After adjusting all seventy-three numbers, the network runs the training data again. It repeats this process thousands of times. Each pass is called an epoch. With enough epochs, the error shrinks to near zero and the network has learned the pattern.
The Training Loop in PHP
Here is the core of the training code. It runs the forward pass, computes the error, and updates every weight and bias for each training example, repeated for five thousand epochs.
$learningRate = 0.5;$epochs = 5000;for ($epoch = 0; $epoch < $epochs; $epoch++) { foreach ($trainingData as $sample) { [$inputs, $target] = $sample; // Forward pass: hidden layer $hidden = []; for ($i = 0; $i < $hiddenSize; $i++) { $sum = $b1[$i]; for ($j = 0; $j < $inputSize; $j++) { $sum += $W1[$i][$j] * $inputs[$j]; } $hidden[$i] = sigmoid($sum); } // Forward pass: output layer $output = sigmoid($b2 + dot($W2[0], $hidden)); // Backward pass: output error $outputDelta = ($target - $output) * sigmoidDerivative($output); // Backward pass: hidden error $hiddenDelta = []; for ($i = 0; $i < $hiddenSize; $i++) { $hiddenDelta[$i] = $outputDelta * $W2[0][$i] * sigmoidDerivative($hidden[$i]); } // Update output weights and bias for ($i = 0; $i < $hiddenSize; $i++) { $W2[0][$i] += $learningRate * $outputDelta * $hidden[$i]; } $b2 += $learningRate * $outputDelta; // Update hidden weights and biases for ($i = 0; $i < $hiddenSize; $i++) { for ($j = 0; $j < $inputSize; $j++) { $W1[$i][$j] += $learningRate * $hiddenDelta[$i] * $inputs[$j]; } $b1[$i] += $learningRate * $hiddenDelta[$i]; } }} For each training example, the code computes a prediction, measures the error, works backward through the network to see how much each weight contributed to that error, and nudges every weight slightly in the direction that reduces it. Repeating this five thousand times over six examples is enough to learn the pattern.
Step 5: Saving the Model
Once training is done, the weights and biases are saved to a JSON file. The file contains five fields:
- vocabulary, the list of ten words
- W1, the input to hidden weights
- b1, the hidden biases
- W2, the hidden to output weights
- b2, the output bias
The JSON file is the entire model. It is a few kilobytes. Nothing else is needed to run predictions.
Step 6: Inference in the Browser
The JSON file is loaded in JavaScript. The same forward propagation is run, this time in the browser. No training happens at inference time. The weights are frozen.
Here is the complete inference function:
function predict(text, model) { // Step 1: turn the text into a 0/1 vector const words = text.toLowerCase().match(/[a-z]+/g) || []; const inputs = model.vocabulary.map(w => words.includes(w) ? 1 : 0); // Step 2: hidden layer const hidden = model.b1.map((bias, i) => { let sum = bias; for (let j = 0; j < inputs.length; j++) { sum += model.W1[i][j] * inputs[j]; } return 1 / (1 + Math.exp(-sum)); }); // Step 3: output layer let outSum = model.b2[0]; for (let i = 0; i < hidden.length; i++) { outSum += model.W2[0][i] * hidden[i]; } const score = 1 / (1 + Math.exp(-outSum)); // Step 4: verdict return { score, label: score > 0.5 ? "SPAM" : "NOT SPAM" };} The whole inference step is eighteen lines and runs in under a millisecond. For the message "click here to win free money," the input vector is [1, 1, 1, 1, 0, 0, 0, 0, 0, 0]. The network multiplies, adds, applies sigmoid, and produces a score of roughly 0.98. That score is above 0.5, so the message is classified as spam.
For "hello, let's discuss the project report," the input vector is [0, 0, 0, 0, 0, 0, 1, 1, 1, 0]. The network produces a score near 0.01. That is below 0.5, so the message is classified as not spam.
What the Hidden Neurons Learned
The most interesting part of this exercise is what the hidden neurons end up detecting, without being told.
Hidden neuron 5 developed strong positive weights for "win," "click," and "offer." It fires hard on marketing language. Its output weight in W2 is the largest positive number in the model, at roughly 3.76.
Hidden neuron 4 developed strong positive weights for "project" and "report." It fires hard on work language. Its output weight is the largest negative number, at roughly -3.83.
Nobody told the network that "project" and "win" belong to different categories. It discovered that on its own, by adjusting numbers to reduce error.
That is the essence of what people mean when they say a neural network "learns." It is not learning in the human sense. It is finding numerical patterns that minimize error. But the result is a system that behaves as if it understood something.
What This Does Not Do
A six-example, ten-word network is not a spam filter. It is a demonstration. Real spam filters train on millions of messages, use thousands of features, and combine the text signal with metadata such as sender reputation, link analysis, and header inspection.
A single typo defeats this network. "FREE M0NEY" becomes [0, 0, 0, 0, 0, 0, 0, 0, 0, 0] because "m0ney" is not in the vocabulary. The network has no way to handle variations it has not seen.
That limitation is not a flaw of the approach. It is a description of how narrow AI works. It only knows what it was trained on.
Why This Matters in 2026
The current generation of AI systems is built on the same principles described here, scaled up. GPT models have billions of parameters instead of seventy-three. They are trained on trillions of tokens instead of six short messages. They use the same forward propagation, the same backpropagation, and the same sigmoid or its modern equivalents.
The specific architecture has evolved. Modern systems use the Transformer, which adds self-attention to the basic feedforward structure. But the arithmetic at each layer is still weights, biases, and non-linear functions. The scale and the architecture differ. The foundation does not.
Understanding the small version makes the large version less mysterious. It also makes the limits of AI clearer. A system that learns from data can only reflect patterns in that data. It cannot reason about things it has never seen. It cannot tell truth from falsehood. It cannot decide what is important. Those are still human problems.
Part 2: The Strategic Decision
For executives, procurement teams, and business owners. If you are here for the code walkthrough, return to Part 1 above.
The next sections explain why Wholesale Dito Store keeps AI out of the customer data path and the calculation path, even though the mechanic documented in Part 1 is one the technical co-founder worked with directly between 2016 and 2019.
Why We Still Exclude AI From Our Customer Path
The technical co-founder built narrow AI systems between 2016 and 2019, while the interest was active. That work stopped before 2020, and the fascination faded. The demonstration in this article is a return to the mechanic for one purpose, not a continuation of an ongoing practice.
And yet Wholesale Dito Store does not run AI in the customer data path, and does not run AI in the calculation path. The reason is not unfamiliarity. It is the opposite.
An AI model is a dependency. It behaves differently on different runs. For pricing, tax, and financial calculations, that variance is not acceptable because the output is the result. And for customer data, an AI model means the data leaves the operator's infrastructure and enters the vendor's, even when the vendor publishes a policy against training on customer data.
Wholesale Dito Store operates its own wholesale business on its own stack. Pricing, RFQs, purchase orders, and customer information stay inside infrastructure the company controls. The Ask W tools run deterministic arithmetic in the browser. The RFQ automation uses pattern matching, not a language model.
The full reasoning is documented in Why We Built a Zero-Dependency Stack: The Engineering Decisions Behind the Wholesale Dito Store Platform.
How to Reproduce This
The training script is a single PHP file. It defines the vocabulary, the training examples, the network architecture, and the training loop. It writes the trained weights to weights.json.
The inference script is a single JavaScript file. It loads weights.json, runs forward propagation, and returns a label and a confidence score.
No dependencies. No package manager. No build step. A local PHP server is enough to run both.
A live demo is available at wholesaledito.store/ask-w/narrow-ai-demo. It runs entirely in the browser. No data is stored or transmitted.
If you want to extend it, add more vocabulary words, more training examples, or additional hidden neurons. The same code structure will hold. The accuracy will improve with the data, up to a point. Beyond that point, you would need a different architecture, which is where the standard machine learning libraries start to matter.
Frequently Asked Questions
What is a narrow AI?
A narrow AI is a system that performs one specific task. It does not generalize to other tasks. A spam classifier is a narrow AI. A chess engine is a narrow AI. A large language model is also a narrow AI, despite its breadth, because it does one thing: predict the next token.
How does a neural network learn?
It adjusts its weights and biases to reduce the error between its predictions and the correct answers. The adjustment is done through backpropagation, which uses the chain rule to determine how much each number contributed to the error.
Can AI be built with PHP and JavaScript?
Yes, for small networks. PHP handles the training loop, and JavaScript handles inference in the browser. For larger networks or more complex tasks, Python and its machine learning libraries are the practical choice.
What is the difference between training and inference?
Training is the process of adjusting weights. Inference is the process of using trained weights to make a prediction. Training happens once, offline. Inference happens every time the model is used.
Why use sigmoid as the activation function?
Sigmoid squashes any number into a range between 0 and 1, which is useful for binary classification. Modern networks often use ReLU because it is faster to compute and avoids the vanishing gradient problem in deep networks.
What is backpropagation?
Backpropagation is the algorithm that calculates how much each weight and bias in the network contributed to the error, so they can be adjusted. It is an application of the chain rule from calculus.
What is an epoch?
An epoch is one full pass through the training data. Training usually involves many epochs. In this example, the network is trained for five thousand epochs.
Can this approach scale to real spam filtering?
Not directly. Real spam filters use much larger vocabularies, much more training data, and additional signals such as sender reputation and link analysis. The mechanics are the same, but the scale is different.
What are the limits of narrow AI?
A narrow AI cannot reason about anything outside its training data. It cannot handle inputs that are meaningfully different from what it saw during training. It cannot explain its decisions in human terms, though the internal computation can be inspected.
Did Wholesale Dito Store use AI to build this?
No. The training script is standard backpropagation written by hand. The inference script runs the same arithmetic in the browser. There is no external AI service involved, and no generative model in the pipeline.
Where can I try the demo?
The live demo is at wholesaledito.store/ask-w/narrow-ai-demo. The demo runs entirely in the browser. No data is stored or transmitted.
Summary
A narrow AI is a system that does one thing. A neural network is a stack of matrix multiplications with non-linear functions between the layers. Training is the process of adjusting the network's weights and biases to reduce error. Inference is the process of using the trained network to make a prediction.
A spam classifier with ten inputs, six hidden neurons, and one output neuron can be trained in PHP and run in JavaScript. It is not a production spam filter, but it is a complete, working example of the mechanics behind every neural network in use today.
Outro
Published by The Sniffer, the strategic insights blog of Wholesale Dito Store. This article is provided as a public reference for developers, business owners, and procurement teams in the Philippines.
The company's history is documented on the Wholesale Dito Store company profile. The company's policy on AI, including the exclusion of AI from the calculation path and the customer data path, is documented in Why We Built a Zero-Dependency Stack.
Wholesale Dito Store is operated by Clickerwayne Zelle Solutions Inc, Forest Drive St., corner Country Drive, Country Homes, Biñan, Laguna 4024, Philippines. Questions can be sent to customercare@wholesaledito.store.