Note 1
Foundations and Intuitions

This note covers basic concepts as well as intuitions through the introduction of Large Language Models.

1.1 Fundamental Terms and Concepts

This section is an introduction to basic terms and concepts that will be used throughout my notes. That being said, more advanced concepts will be introduced in latter notes. Please feel free to skip over this section if you already know the bold words.

Definition 1.1.1

Machine learning is a subset of artificial intelligence that learns from data. Deep learning is a subset of machine learning that utilizes neural networks.

The primary purpose of ML is to make AI “smarter”. For instance, calculators are “dumb” in a sense that it requires an input and can only do certain tasks such as outputting calculated input. However, we can make it “smarter” by teaching data. We can clearly see that WolframAlpha is smarter than a normal scientific calculator. (Please note that this is just an analogy not an example because WolframAlpha is a knowledge engine, not a machine learning model. But I mean you get the point.)

On the high level, you can think of an AI model as a giant function with billions of array of numbers, called parameters.

Definition 1.1.2

Parameters are the internal values in a model composed with weights and biases that are learned by the model. After the parameters are finalized with the training phase, the models generate predictions through the process called inference.

Indeed, models are prediction machines based on the trained parameters. One fun fact here is that it generally costs way less to run inference than to train the model. More in-depth explanation will be available in latter notes where we discuss neural nets.

1.1.1 ML Paradigms

In ML, there are three primary learning paradigms: supervised learning, unsupervised learning, and reinforcement learning. First, what does it mean to “learn” and “train”?

Definition 1.1.3

Machine learning utilizes the process of training to learn from data with algorithms. A common example would be dog and cat recognition from picture with labeled data.

There are multiple ways in which models are trained. One of the most used method out of the three paradigms is supervised learning.

Definition 1.1.4

When a model is trained with supervised learning, the model utilizes labeled dataset, or examples of input-output pairs, to learn.

With supervised learning, there are two primary predictions that models make.

Definition 1.1.5

Classification is a prediction where the label of an input data is being predicted. Regression is utilized to predict continuous values.

Classic examples of classification would be categorizing dog/cat pictures and email spam filtering. An example of regression would be the house price over years.

The second most common learning paradigm is unsupervised learning.

Definition 1.1.6

Unsupervised learning is a process of training a model with unlabeled data.

How is it possible for a model to learn without labeled data? One primary example algorithm is \(k\)-means clustering, where the model categorizes the data with relevance and similarity. To cluster high dimensional data, we often perform dimensionality reduction to reduce the dimension for clustering with the trust in manifold hypothesis. I won’t go to deep into unsupervised learning and \(k\)-means clustering in this note, but you can check out this awesome visualization of \(k\)-means clustering.

The third most used learning paradigm is reinforcement learning.

Definition 1.1.7

Reinforcement learning is a process of training models through interactions with environment and reward system.

Personally, reinforcement learning is more complicated compared to the two paradigms introduced above. On the high level, an agent will take an action in an environment, where it returns rewards based on the action. The agent will then update its policies with the provided reward. The process is continued until the “adequate” amount is reached.

Training a model can be understood as writing a best fit mathematical function to training datasets that captures the holistic trend. There are few terms that describe the function.

Definition 1.1.8

A model is overfit if it learns too much from the data that it does not effectively capture that holistic pattern but small noises. On the other hand, an underfitting model is formed when it does not learn effectively from the data.

On the high level, our goal is to generalize and capture the general pattern and prevent overfitting or underfitting. We normally would stop training a model when it does not appear to improve, or is degrading, but how do we assess it?

Definition 1.1.9

To prevent overfitting and to evaluate the model during training, validation is utilized with unseen data. Unlike validation, test sets are used at the end of the training to evaluate the finalized model.

We have discussed the fundamental terms and concepts that would repeatedly appear in the following notes. Now let’s build our high level intuitions on ML with the introduction to LLMs!