0.1 Building Vanilla NN From Scratch

For our toy examples, we will create two things. The first one is technically not a neural network as it does not contain conventional neurons and structures. The second one is an actual neural net that is very simple. In this section, I will share the entire code first, then provide explanations.

0.1.1 Simple Stochastic Gradient Descent-Based
Learning Engine

What we will make now is a simple stochastic gradient descent-based learning engine. This isn’t technically a neural network, as you will see why in the code, but I decided to include it here as making similar examples really helped me understand backprop with code.

Our task with this learning engine is simple. We are going to predict \(e^x\) for limited range with concatenated Maclaurin Series. If you recall from your calculus courses, the following equation holds by Taylor series. \[ e^x = \sum _{n = 0}^\infty \frac {1}{n!} x^n = 1 + x + \frac {x^2}{2!} + \frac {x^3}{3!} + \cdots \] Here, for limited \(x\), we will be approximating \(e^x\) with the following polynomial where \(b\) is the bias and \(w_i\) are the weights. \[ e^x \approx b + w_1x + w_2x^2 + w_3x^3 + w_4x^4 + w_5x^5 + w_6x^6 + w_7x^7 + w_8x^8 \] We will get back to this on why our \(x\) should be limited, but let’s see our structure first.

We will be “training” the engine with \(200\) examples from \(-1.00\) to \(0.99\) inclusive. Then, we will be testing it with \(50\) validations from \(1.00\) to \(1.49\) inclusive. This backprop machine is stochastic as it will adhere to the following structure.

[Picture]

Figure 1: Overview of the Simple Learning Engine

As you can see, we will be using MSE for our loss function, and we do not need activation functions like sigmoid and tanh because we don’t need one in our simple example. Okay now here is the breakdown of the entire notebook available here.

Detailed Explanation

Now let’s take a look at individual cells.

1import math 
2import random 
3import matplotlib.pyplot as plt 
4%matplotlib inline

Here, I tried to keep it as minimal as possible. Instead of NumPy, I decided to use the math module for all math stuffs like math.exp(). We need random as it would make our lives easier than to manually set the eight weights and the bias. Matplotlib of course since we want to see how it learns for limited range of input.

1train_x = [i/100 for i in range(-100, 100)] 
2train_y = [math.exp(i/100) for i in range(-100, 100)] 
3test_x = [i/100 for i in range(100, 150)] 
4test_y = [math.exp(i/100) for i in range(100, 150)]

Setting up the dataset for our engine is also quite simple. Instead of writing every single number, I just decided to loop through them and store them in an array, and that is train_x. The values for train_y is also quite straightforward as it is \(e^x\) for corresponding \(x\). The same goes for the test set. I tried to keep it \(80\%\) and \(20\%\) for training and validation, but that is not necessary and it varies by models. However, I decided to use that here because that’s pretty much what most people do when they first train. Now let’s get into the fun stuffs.

1class SGDBLE: 
2    def __init__(self): 
3        random.seed(3141592) # So that we get the same result 
4        self.ws = [random.uniform(-1, 1) for _ in range(8)] 
5        self.b = random.random()

Now I decided to store the core functions in the class SGDBLE as we will be using self to store attributes and objects. Here we set the specific seed so that you will also get the same result when you run the code in somewhere. Next we set weights and biases and stored them in self with our set random values.

1def forward(self, x): 
2    self.ypred = sum(w * x**k for w, k in zip(self.ws,range(1, 9))) + self.b 
3    return self.ypred 
4 
5def loss(self, y, ypred): 
6    self.loss_val = 0.5 * (y - ypred)**2 
7    return self.loss_val

This is where we set some core functions. Forward is forward pass and it is somewhat straight forward as that is how we defined it. For the loss function, we are using MSE, i.e. \(\frac {1}{2} (y - \hat {y})^2\).

1def grads(self, x, y, ypred): 
2    base_der = ypred - y # i.e. dL / dypred 
3    self.grads_w = [base_der * x**i for i in range(1, 9)] 
4    self.grad_b = base_der 
5 
6def update(self, lr): 
7    for i in range(8): 
8        self.ws[i] -= lr * self.grads_w[i] 
9    self.b -= lr * self.grad_b

This is the core part of our engine. Let’s briefly look at the math and write the loop out for some to see what’s actually going on. Consider the following equation by the chain rule. \[ \frac {dL}{dw_i} = \frac {dL}{dy} \cdot \frac {dy}{dw_i} \] Notice that \(\frac {dL}{dy} = \frac {d}{dy} (\frac {1}{2} (y - \hat {y})^2) = y - \hat {y}\). This is base_der. Continuing with the chain rule, we can write the weights and the bias as the following. \begin{align*} \frac {dL}{dw_1} &= \frac {dL}{dy} \cdot \frac {dy}{dw_1} = \text {base\_der} \cdot x \\ \frac {dL}{dw_2} &= \frac {dL}{dy} \cdot \frac {dy}{dw_2} = \text {base\_der} \cdot x^2 \\ \frac {dL}{db} &= \frac {dL}{dy} \cdot \frac {dy}{db} = \text {base\_der} \cdot 1 \end{align*}

Now the update function. Recall \(w_2 = w_1 - \eta \frac {dL}{dw_1}\) and \(b_2 = b_1 - \eta \frac {dL}{db_1}\). That is exactly what we did, but for weights we have to loop through each element.

1def train(self, epochs, lr): 
2    losses = [] 
3    for epoch in range(epochs): 
4        loss_total = 0 
5        for i in range(len(train_x)): 
6            ypred = self.forward(train_x[i]) 
7            loss_total += self.loss(train_y[i], ypred) 
8            self.grads(train_x[i], train_y[i], ypred) 
9            self.update(lr) 
10        loss_avg = loss_total / len(train_x) 
11        losses.append(loss_avg) 
12 
13        if (epoch+1) % 100 == 0: 
14            print(f"Epoch: {epoch+1}/{epochs}, Loss: {loss_avg}") 
15 
16    return losses

This is the training loop. Of course we don’t want to stop at a single epoch especially in stochastic setting. You can ignore the losses in the beginning as we only use them to graph. For each epoch, we want to set the total loss to \(0\) as they will accumulate. Then, for all \(x\) values in our training set, we will predict, calculate the loss, backprop, an update as shown. Note that we have to set epoch and lr, which is the learning rate. Now we will take the average of the total loss and print it every \(100\) epoch to see how well our engine is learning.

1def test(self): 
2    loss_total = 0 
3    for i in range(0, len(test_x), 2): 
4        ypred1 = self.forward(test_x[i]) 
5        loss_total += self.loss(test_y[i], ypred1) 
6 
7        if i + 1  len(test_x): 
8            ypred2 = self.forward(test_x[i + 1]) 
9            loss_total += self.loss(test_y[i + 1], ypred2) 
10 
11            print(f"Test {i+1}                      Test {i+2}") 
12            print(f"y:     {test_y[i]}      y:     {test_y[i+1]}") 
13            print(f"ypred: {ypred1}     ypred: {ypred2}") 
14            print() 
15        else: 
16            print(f"Test {i+1}") 
17            print(f"y:     {test_y[i]}") 
18            print(f"ypred: {ypred1}") 
19 
20    loss_total /= len(test_x) 
21 
22def print_params(self): 
23    coeff = [1 / math.factorial(i+1) for i in range(8)] 
24    for i in range(8): 
25        print(f"w{i+1}: {self.ws[i]}, Expected: {coeff[i]}") 
26    print(f"b: {self.b}, Expected: 1") 
27 
28model = SGDBLE()

The test loop is also straight forward. It predicts, calculates the loss, and print them as shown. Also, because we know the coefficients for the Maclaurin Series, I decided to see if the parameters match with the actual coefficients. Finally, the last line is to be able use the model.

1def graph_loss(losses): 
2    plt.plot(losses) 
3    plt.xlabel("Epoch") 
4    plt.ylabel("Loss") 
5    plt.show() 
6 
7def graph_pred(): 
8    xs = [i / 10 for i in range(-100, 200)] 
9    true = [math.exp(x) for x in xs] 
10    pred = [model.forward(x) for x in xs] 
11 
12    plt.plot(xs, true, label="e^x") 
13    plt.plot(xs, pred, label="Prediction") 
14    plt.yscale("log") 
15    plt.legend() 
16    plt.show()

This part is plotting with Matplotlib. Please refer to the second note if you are confused! One thing though yscale is necessary as exponential function will get too large to show.

1losses = model.train(1000, 0.005) 
2graph_loss(losses) 
3model.print_params() 
4graph_pred()

This part is the actual training part. We set the epoch to \(1000\) and learning rate to \(0.005\). This learning rate might be too small for the start, but I realized that this gives slightly better results than \(0.001\) in this setting. Then we graph and show loss, parameters, and predictions.

1model.test()

This part is for testing. Now that we have seen the entire code and explanation for it, let’s now analyze our results.

Analyzing the Results

First, let’s take a look at the first result.

1losses = model.train(1000, 0.005) 
2graph_loss(losses) 
3model.print_params() 
4graph_pred()
Epoch: 100/1000, Loss: 0.00013234985876534958 
Epoch: 200/1000, Loss: 3.427320321762344e-05 
Epoch: 300/1000, Loss: 2.4492707827970274e-05 
Epoch: 400/1000, Loss: 2.2669862503550263e-05 
Epoch: 500/1000, Loss: 2.172205317266906e-05 
Epoch: 600/1000, Loss: 2.0984248749352728e-05 
Epoch: 700/1000, Loss: 2.036538265870468e-05 
Epoch: 800/1000, Loss: 1.9834658817269138e-05 
Epoch: 900/1000, Loss: 1.937270927247297e-05 
Epoch: 1000/1000, Loss: 1.8965386261105253e-05

PIC

w1: 0.9444454598562942, Expected: 1.0 
w2: 0.49986306356857685, Expected: 0.5 
w3: 0.620431907367758, Expected: 0.16666666666666666 
w4: 0.21618569290559692, Expected: 0.041666666666666664 
w5: -0.9321256839917001, Expected: 0.008333333333333333 
w6: -0.48709437200615985, Expected: 0.001388888888888889 
w7: 0.5594822789317027, Expected: 0.0001984126984126984 
w8: 0.33002989302732855, Expected: 2.48015873015873e-05 
b: 0.9983077779493585, Expected: 1

PIC

This is what we got after the first training. The loss looks pretty good actually. As you can see for each \(100^\text {th}\) epoch, the loss is decreasing well. From the graph, you can see the sudden drop in losses; however, we can’t necessarily say that this is a grokking phenomena as it might just be model understanding the pattern from underfitting.

Looking at the weights, some are good and some are bad. We will get back to this later, but this is also one of the reason why we cannot scale this engine. Similarly, we can see that our polynomial fits very well for the small range, but not well as the input increases.

1model.test()
Test 1                    Test 2 
y:     2.718281828459045    y:     2.7456010150169163 
ypred: 2.7495260176087557   ypred: 2.7867851769675642 
 
Test 3                    Test 4 
y:     2.7731947639642978   y:     2.801065834699079 
ypred: 2.8257757734623574   ypred: 2.8666322590606814 
 
Test 5                    Test 6 
y:     2.82921701435156     y:     2.857651118063164 
ypred: 2.909497162739583    ypred: 2.954521430346852 
 
Test 7                    Test 8 
y:     2.8863709892679585   y:     2.915379499976997 
ypred: 3.001864774074873    ypred: 3.0516960317129964 
 
Test 9                    Test 10 
y:     2.944679551065524    y:     2.9742740725630656 
ypred: 3.1041935358457007   ypred: 3.159545493165112 
 
Test 11                   Test 12 
y:     3.0041660239464334   y:     3.034358394435676 
ypred: 3.2179503740678106   ypred: 3.2796173127071624 
 
Test 13                   Test 14 
y:     3.0648542032930024   y:     3.095656500124711 
ypred: 3.3447665176737535   ypred: 3.4136296934778425 
 
Test 15                   Test 16 
y:     3.1267683651861553   y:     3.158192909689767 
ypred: 3.4864504730090555   ypred: 3.5634848611499024 
... 
Test 45                   Test 46 
y:     4.220695816996552    y:     4.263114515168817 
ypred: 9.347554553339133    ypred: 9.752958544358306 
 
Test 47                   Test 48 
y:     4.305959528345206    y:     4.349235141062741 
ypred: 10.179646340417854   ypred: 10.628580569357089 
 
Test 49                   Test 50 
y:     4.392945680918757    y:     4.437095519003664 
ypred: 11.100758361572902   ypred: 11.597212287422602

For the training set, I only listed some of them as you can refer to the previous pages. Notice that for first few tests, the predictions are not bad. However, for Test 5, the engine does not seem be effective and seems lost from Test 10. We could actually run our training loop multiple times to see what happens, and I will leave this to the readers. However, the engine does not seem to learn well.

Here is the question. Is the engine overfitting right now? If you said no, well... that is correct! Technically, we can see it as overfitting, but what I think is more accurate is to say that this is the limitation of the scalability of this engine and our method. When predicting, we only used nine terms. From the parameters, some were very close to the expectation and others were very bad. This is because we are using a polynomial with finite term to predict, whereas the real Maclaurin series uses infinite sum. Therefore, it is techinically an overfitting to certain point in a way that our engine tries to fit our polynomial in limited range as shown in the graph, but this too could improve if we have the computational ability to use infinite sum. Therefore, it may be more accurate to say that it is the limitation of the engine and scalability that causes high loss in testing.

We just completed our first model! Although it is very limited in terms of scalability and usability, I hope this example was very helpful in using backpropagation in code. For our next example, we will build an actual simple neural net to predict a nonlinear function.

0.1.2 Vanilla Neural Network

Now that we have created a simple backprop engine, let’s increase the number of hidden layers to predict \(x_1 x_2\) for inputs \(x_1\) and \(x_2\). I thought it would be great to introduce nonlinearity to our problem, and it works! Let’s started with detailed explanations of the entire notebook available here.

Detailed Explanation

Now let’s take a look at each block of the code to see their role in our vanilla neural net.

1import math 
2import random 
3import matplotlib.pyplot as plt 
4%matplotlib inline

Here, I imported essential modules and libraries to do math, get a randomized value, and plot the loss.

1# Predicting x1 * x2 
2def tanh(x): 
3    return (math.exp(x) - math.exp(-x)) / (math.exp(x) + math.exp(-x)) 
4 
5def tanh_der(x): 
6    return 1 - tanh(x)**2 
7 
8def loss(y, ypred): 
9    return 0.5 * (y - ypred)**2

For this problem, I decided to use tanh for our activation function as sigmoid squishes the values to 0 and 1, which is not ideal for regression problem like ours and ReLU was not effective when I tried. Since \(\tanh (x) = \frac {e^x - e^{-x}}{e^x + e^{-x}}\) and \(\tanh '(x) = 1 - \tanh ^2(x)\), we did the same. For our loss function, I decided to use MSE for simplicity again.

1# Data 
2data_x = [[i/10, j/10] for i in range(-10, 11) for j in range(-10, 11)] 
3data_y = [(x1 * x2) for x1, x2 in data_x] 
4 
5data = list(zip(data_x, data_y)) 
6random.seed(3141592) 
7random.shuffle(data) 
8 
9split = int(0.8 * len(data)) 
10train_data = data[:split] 
11test_data = data[split:] 
12 
13train_x = [x for x, y in train_data] 
14train_y = [y for x, y in train_data] 
15test_x = [x for x, y in test_data] 
16test_y = [y for x, y in test_data]

This part is for making our datasets. I decided to divide \(i\) and \(j\) by \(10\) because if I don’t, math overflow error will be resulted due to \(\tanh \) function.

This time, I decided to shuffle data in a way that we will get the same result when you try running this, and split the data into \(80\%\) for training and \(20\%\) for validating.

1class MLP: 
2    def __init__(self, layers): 
3        self.ws = [] 
4        self.bs = [] 
5 
6        random.seed(3141592) 
7        for i in range(len(layers) - 1): 
8            layer_in = layers[i] 
9            layer_out = layers[i + 1] 
10            layer_w = [[random.uniform(-1, 1) for _ in range(layer_in)] 
11                       for _ in range(layer_out)] 
12            layer_b = [random.uniform(-1, 1) for _ in range(layer_out)] 
13 
14            self.ws.append(layer_w) 
15            self.bs.append(layer_b)

Again for our example, I decided to store our core functions in a class since we want to reuse and store our attributes. When this class is first called, we will create an empty array and for ws and bs and store them in self. Then, for the randomized value, we will loop through each layers to store a randomized value between \(-1\) and \(1\) for weights and biases. Notice that unlike the bias which only loop through the neurons in the output layer, we have to loop through the previous layer too for weights since we have different weights for each previous neuron for each output neuron. We will then append, or store, our layer to ws and bs.

1def forward(self, x): 
2    self.act = [x] 
3    self.zs = [] 
4 
5    # Layers 
6    for i in range(len(self.ws)): 
7        prev = self.act[-1] 
8        pre_layer = [] 
9        act_layer = [] 
10 
11        # Neurons 
12        for j in range(len(self.ws[i])): 
13            pre = self.bs[i][j] 
14            for k in range(len(self.ws[i][j])): 
15                pre += self.ws[i][j][k] * prev[k] 
16 
17            pre_layer.append(pre) 
18 
19            if i != len(self.ws) - 1: 
20                act_layer.append(tanh(pre)) 
21            else: 
22                act_layer.append(pre) 
23 
24        self.zs.append(pre_layer) 
25        self.act.append(act_layer) 
26 
27    return self.act[-1][0]

This is the forward function. Here, I decided to create both self.act and self.zs since we need both of them for backprop function. As the name suggests, self.act is the value after \(\tanh \) is applied to self.zs. Moreover, I stored the input in self.act in the beginning to start using them in the first hidden layer.

Continuing, we create arrays for each layer for preactivation and activation. Then, we will continue by looping through the neurons within each layer. Our preactivation will start from the bias, and by following the definition, we will add the products of weights and the previous activation. For the if-else condition, I decided to do this because we don’t want \(\tanh \) for our final output layer. Then, we will store the layers in self.act and self.zs for latter use.

1def backprop(self, y): 
2    self.grad_w = [[[0.0 for _ in range(len(self.ws[i][j]))] 
3                    for j in range(len(self.ws[i]))] 
4                   for i in range(len(self.ws))] 
5    self.grad_b = [[0.0 for _ in range(len(self.bs[i]))] 
6                   for i in range(len(self.bs))] 
7    ypred = self.act[-1][0] 
8    delta = [0] * len(self.ws) 
9    delta[-1] = [ypred - y] 
10 
11    # range(start, stop, step) 
12    for k in range(len(self.ws) - 2, -1, -1): 
13        delta[k] = [] 
14        for i in range(len(self.ws[k])): 
15            dlda = 0.0 
16            for j in range(len(self.ws[k + 1])): 
17                dlda += self.ws[k + 1][j][i] * delta[k + 1][j] 
18            delta[k].append(dlda * tanh_der(self.zs[k][i])) 
19 
20    for i in range(len(self.ws)): 
21        prev = self.act[i] 
22        for j in range(len(self.ws[i])): 
23            self.grad_b[i][j] = delta[i][j] 
24            for k in range(len(prev)): 
25                self.grad_w[i][j][k] = delta[i][j] * prev[k]

Ah the backprop. I think this is the hardest part from our code. Assuming that you have read the previous section, let’s interpret the code with the corresponding math equations.

Recall although we do not calcuate the loss per individual parameter, we do have to calculate the gradients. To start, I decided to put the initial values as \(0\). To store them, we can do the same just as we randomized the initial weights and biases because ultimately, each gradient will have a matching parameter. Continuing, ypred is the prediction, and that is the first element from the last activation layer. This will vary when we have numerous outputs, but since this is a simple regression problem and we only have one output, this is good. delta here is the delta term for our hidden layers. Let’s start by setting them as \(0\) at first and ypred - y for the last layer.

Here comes the fun part. In Python, the range function can be used as range(start, stop,step). Here, we start from the last hidden layer and we subtract \(2\) due to indices. Then, we will stop when we reach \(-1\) by going backward by \(1\). We will loop through them for our delta terms. Recall the formula that we derived below. \[ \delta _i^{[l]} = \frac {\partial L}{\partial a_i^{[l]}} \frac {\partial a_i^{[l]}}{\partial z_i^{[l]}} = g'\left (z_i^{[l]}\right ) \sum _j w_{ji}^{[l]} \delta _j^{[l+1]} \] As the formula suggests, we will add the weights and the corresponding following delta term with the for loop for \(\frac {\partial L}{\partial a_i^{[l]}}\). Then, we will mulitply it by the derivative of the actiavtion function, which is \(\tanh \) for our example. Finally, we can store our delta terms that are calculuated. Next, we can now use them to calculate the gradients. \begin{align*} \frac {\partial L}{\partial w_{ij}^{[l]}} &= \delta _i^{[l]} a_j^{[l-1]} \\ \frac {\partial L}{\partial b_i^{[l]}} &= \delta _i^{[l]} \frac {\partial z_i^{[l]}}{\partial b_i^{[l]}} = \delta _i^{[l]} \end{align*}

As the formula suggests, for each parameters, will calculate the gradients again with for the loop. This is it for the backprop!

1def update(self, lr): 
2    for i in range(len(self.ws)): 
3        for j in range(len(self.ws[i])): 
4            for k in range(len(self.ws[i][j])): 
5                self.ws[i][j][k] -= lr * self.grad_w[i][j][k] 
6            self.bs[i][j] -= lr * self.grad_b[i][j] 
7 
8def train(self, train_x, train_y, epochs, lr): 
9    losses = [] 
10    for epoch in range(epochs): 
11        total_loss = 0 
12 
13        for x, y in zip(train_x, train_y): 
14            ypred = self.forward(x) 
15            total_loss += loss(y, ypred) 
16            self.backprop(y) 
17            self.update(lr) 
18 
19        average_loss = total_loss / len(train_x) 
20        losses.append(average_loss) 
21 
22        if (epoch + 1) % 100 == 0: 
23            print(f"Epoch: {epoch+1}/{epochs}, Loss: {average_loss}") 
24 
25    return losses

Now we need to update. This part is similar with our previous example, but for each parameter, we will update it by subracting the product of the corresponding gradient and the learning rate. The training part is again very similar to our previous example. We will begin our loss with \(0\), the we will predict, accumulate the loss, backprop, update, and print every hundred epoch.

1    def show_params(self): 
2        for i, (w, b) in enumerate(zip(self.ws, self.bs)): 
3            layer = "OUTPUT" if i == len(self.ws) - 1 else f"LAYER {i+1}" 
4            print(layer) 
5            for j, neuron_w in enumerate(w): 
6                weights_str = ", ".join(f"{x:.4f}" for x in neuron_w) 
7                print(f"  Neuron {j+1}:") 
8                print(f"    Weights: [{weights_str}]") 
9                print(f"    Bias = {b[j]:.4f}") 
10            print() 
11 
12mlp = MLP([2, 8, 8, 1])

This part is for printing the parameters, and as you can see, it is pretty straight forward. I ended up not using it because I don’t think it is that effective at this point, but why not just have it. Now we create an instance for the MLP class with two hidden layers with eight neurons each.

1losses = mlp.train(train_x, train_y, epochs=2000, lr=0.01) 
2print() 
3 
4for x, y in zip(train_x[:10], train_y[:10]): 
5    print(x, y) 
6    print(f"     ypred: {mlp.forward(x)}") 
7 
8plt.plot(losses) 
9plt.xlabel("epoch") 
10plt.ylabel("loss") 
11plt.show() 
12 
13#mlp.show_params()

We will train with the learning rate of \(0.01\) and \(2000\) epochs. We then print the first ten p9redictions and plot the loss.

1test_loss = 0 
2for x, y in zip(test_x, test_y): 
3    ypred = mlp.forward(x) 
4    test_loss += loss(y, ypred) 
5 
6for x, y in zip(test_x[:10], test_y[:10]): 
7    ypred = mlp.forward(x) 
8    print(x, y) 
9    print(f"    ypred: {ypred}") 
10 
11print("Test Loss:", test_loss / len(test_x))

This part is for printing the loss and prediciton for testing. That was it for the explanation! Now let’s take a look at the results.

Analyzing the Results

Now that we have looked at the source code, let’s analyze the result to see how our vanilla neural net performs.

1losses = mlp.train(train_x, train_y, epochs=2000, lr=0.01) 
2print() 
3 
4for x, y in zip(train_x[:10], train_y[:10]): 
5    print(x, y) 
6    print(f"     ypred: {mlp.forward(x)}") 
7 
8plt.plot(losses) 
9plt.xlabel("epoch") 
10plt.ylabel("loss") 
11plt.show() 
12 
13#mlp.show_params()
Epoch: 100/2000, Loss: 0.0006795237745380497 
Epoch: 200/2000, Loss: 0.0002780199905385479 
Epoch: 300/2000, Loss: 0.00015780133554953304 
Epoch: 400/2000, Loss: 9.996884725589866e-05 
Epoch: 500/2000, Loss: 6.544552166502369e-05 
Epoch: 600/2000, Loss: 4.373246439583505e-05 
Epoch: 700/2000, Loss: 3.032175895325618e-05 
Epoch: 800/2000, Loss: 2.2295750007135315e-05 
Epoch: 900/2000, Loss: 1.7559747007674843e-05 
Epoch: 1000/2000, Loss: 1.4723606896641237e-05 
Epoch: 1100/2000, Loss: 1.2951440585334302e-05 
Epoch: 1200/2000, Loss: 1.1773900112740833e-05 
Epoch: 1300/2000, Loss: 1.0936643742787753e-05 
Epoch: 1400/2000, Loss: 1.0302935457567813e-05 
Epoch: 1500/2000, Loss: 9.79811628836463e-06 
Epoch: 1600/2000, Loss: 9.379957580934361e-06 
Epoch: 1700/2000, Loss: 9.023360655331247e-06 
Epoch: 1800/2000, Loss: 8.712547584218995e-06 
Epoch: 1900/2000, Loss: 8.437042172730164e-06 
Epoch: 2000/2000, Loss: 8.189547717119376e-06 
 
[1.0, -1.0] -1.0 
     ypred: -0.9914314084230468 
[0.8, -1.0] -0.8 
     ypred: -0.8026975789984329 
[0.8, -0.9] -0.7200000000000001 
     ypred: -0.7229659139930072 
[0.1, -0.8] -0.08000000000000002 
     ypred: -0.07901899113208088 
[0.3, -0.3] -0.09 
     ypred: -0.08931568886644314 
[-0.7, -0.1] 0.06999999999999999 
     ypred: 0.07003157242570901 
[-0.6, -0.5] 0.3 
     ypred: 0.29863205953269367 
[0.7, 0.8] 0.5599999999999999 
     ypred: 0.5613061391881997 
[-0.5, -0.4] 0.2 
     ypred: 0.19629911113434118 
[0.9, 0.8] 0.7200000000000001 
     ypred: 0.7218073084017129

PIC

I would say this is pretty good! The loss is constantly decreasing and the prediction seems to be pretty good. Looking at the graph, the model seemed to have little idea at first, but instantly learned to see the holistic view as suggested by the sudden drop.

1test_loss = 0 
2for x, y in zip(test_x, test_y): 
3    ypred = mlp.forward(x) 
4    test_loss += loss(y, ypred) 
5 
6for x, y in zip(test_x[:10], test_y[:10]): 
7    ypred = mlp.forward(x) 
8    print(x, y) 
9    print(f"    ypred: {ypred}") 
10 
11print("Test Loss:", test_loss / len(test_x))
[-0.9, -0.3] 0.27 
    ypred: 0.27224246623233583 
[0.9, 1.0] 0.9 
    ypred: 0.889162612694173 
[0.1, 0.5] 0.05 
    ypred: 0.050824340568330406 
[1.0, 1.0] 1.0 
    ypred: 0.9737260427905209 
[0.6, -0.8] -0.48 
    ypred: -0.4811987339447982 
[0.6, -0.3] -0.18 
    ypred: -0.17587974783858806 
[0.6, 0.4] 0.24 
    ypred: 0.23577905969192914 
[0.5, -0.8] -0.4 
    ypred: -0.4007379739663608 
[0.5, 0.2] 0.1 
    ypred: 0.0970149401772119 
[0.9, -0.2] -0.18000000000000002 
    ypred: -0.18071247280108016 
Test Loss: 1.1432408362631248e-05

Our test also seems pretty good. The loss seems small because our validation set uses small numbers, but we could see that the prediction somewhat matches the real values!

Now looking back at our model, this is again good for the limited range as suggested by the result above. However, there is a reason for that. Recall that we shuffled the dataset and seletcted \(80\%\) for training and \(20\%\) for testing there. Therefore, for the limited range that the model is trained in, the model did pretty good job. However, when I first tried making dataset by splitting the training and validation as first \(80\%\) and latter \(20\%\), our model did poorly. This is because the model learned the left side well, but it does poorly in predicting the right side since it has no idea what’s going on there with absolutely no training example. However, the good news is that our neural net is scalable unlike the previous example! By increasing the range for datasets, and could see whether the model underfits or overfits.

We now covered two models from scratch! Now that we know what happens in deep neural nets, let’s build our own language models!