The first `nn.Linear` layer maps two input features to eight hidden values. I find the two feature counts a clear way to see what this layer changes.
You’ll place that layer in a model structure that can hold its parameters and connect it to another layer. Next, see how `torch.nn` contributes as you define the classifier.
What torch.nn contributes to a model
torch.nn is PyTorch’s namespace for neural-network modules, layers, containers, activation functions, and loss functions. nn.Module is the base class for a custom model, while torch.nn itself is not a model class.
Assigning a layer to a module attribute registers its parameters with the model, so model.parameters() can pass them to an optimizer. The optimizer lives in torch.optim, a separate namespace covered in this torch.optim guide.
Registration also lets a parent module move its registered parameters and buffers together and include them in its state dictionary. A plain tensor attribute is not a learnable parameter unless you register it as an nn.Parameter.
| Component | What it does | Has learnable state by default |
|---|---|---|
| nn.Module | Base class for a model or reusable model component | Only when you add parameters or child modules |
| nn.Linear | Maps the final input dimension to a new feature dimension | Yes, weight and bias by default |
| torch.nn.functional | Provides function calls such as activations and tensor operations | No module-owned parameters |
| torch.optim | Updates registered parameters using their gradients | Optimizer state is separate from the model |
I chose a small Linear model rather than a CNN because its feature-to-logit shape is visible without introducing image dimensions. Use a module when a layer needs registered parameters, while a torch.nn.functional operation suits a step with no module-owned state, as the PyTorch forum discussion about the two APIs illustrates.
nn.Sequential is itself a module container that calls child modules in order, which is why the example can store both Linear layers under one attribute.
Check the inputs before defining layers
Each row is one example, and the last dimension holds its input features. A Linear layer replaces that feature dimension with out_features, which means the batch dimension stays unchanged.
Linear applies an affine transformation to the final dimension with a weight matrix and optional bias, as its reference defines.
For a two-class classifier, the target stores one integer class ID per example. CrossEntropyLoss compares those IDs with two raw scores per row, so do not apply Softmax before the loss, as its reference specifies.
Here Linear carries six feature rows from (6, 2) to (6, 8) and back to (6, 2), so the batch stays at six while the feature count changes.
import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"device: {device}")
| Tensor | Shape and type | Meaning |
|---|---|---|
| features | (batch, 2), float32 | Two input values for each example |
| targets | (batch,), long | One class ID, 0 or 1, for each example |
| logits | (batch, 2), float32 | Two unnormalized class scores per example |
The device check selects CUDA when available or keeps the example on CPU, and the model, features, and labels must share that device before the forward pass.
Build and train a classifier in four steps
The example uses eight two-feature rows, with two reserved for validation. The same shapes work with larger data as long as each row still has two features.
Step 1: Match the input and class dimensions
The first Linear layer maps two input features to eight hidden values, then the second maps eight values to two class scores that match the class IDs.
Step 2: Register layers inside nn.Module
Classifier inherits from nn.Module and stores an nn.Sequential container as self.layers. The container registers both Linear layers, so model.parameters() includes their weights and biases without a separate parameter list.
Step 3: Connect logits, loss, and optimizer
CrossEntropyLoss takes the model’s raw logits and integer class labels. Each loop clears old gradients, calculates a loss, backpropagates it, and lets stochastic gradient descent adjust the registered parameters.
import torch
from torch import nn
torch.manual_seed(7)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
features = torch.tensor(
[
[-2.0, -1.0],
[-1.0, -2.0],
[-1.5, -0.5],
[-2.0, -2.0],
[1.0, 1.0],
[2.0, 1.0],
[1.0, 2.0],
[1.5, 0.5],
],
dtype=torch.float32,
device=device,
)
labels = torch.tensor([0, 0, 0, 0, 1, 1, 1, 1], dtype=torch.long, device=device)
train_indices = torch.tensor([0, 1, 2, 4, 5, 6], device=device)
validation_indices = torch.tensor([3, 7], device=device)
train_features = features[train_indices]
train_labels = labels[train_indices]
validation_features = features[validation_indices]
validation_labels = labels[validation_indices]
class Classifier(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.Sequential(
nn.Linear(2, 8),
nn.ReLU(),
nn.Linear(8, 2),
)
def forward(self, x):
return self.layers(x)
model = Classifier().to(device)
loss_fn = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
with torch.no_grad():
initial_loss = loss_fn(model(train_features), train_labels).item()
model.train()
for _ in range(120):
optimizer.zero_grad()
logits = model(train_features)
loss = loss_fn(logits, train_labels)
loss.backward()
optimizer.step()
model.eval()
with torch.inference_mode():
train_logits = model(train_features)
validation_logits = model(validation_features)
final_loss = loss_fn(train_logits, train_labels).item()
train_correct = (train_logits.argmax(dim=1) == train_labels).sum().item()
validation_correct = (
validation_logits.argmax(dim=1) == validation_labels
).sum().item()
parameter_count = sum(parameter.numel() for parameter in model.parameters())
print(f"PyTorch: {torch.__version__}")
print(f"device: {device}")
print(f"input shape: {tuple(train_features.shape)}")
print(f"logits shape: {tuple(train_logits.shape)}")
print(f"validation logits shape: {tuple(validation_logits.shape)}")
print(f"trainable parameters: {parameter_count}")
print(f"cross-entropy: {initial_loss:.4f} -> {final_loss:.4f}")
print(f"training accuracy: {train_correct}/{len(train_labels)}")
print(f"validation accuracy: {validation_correct}/{len(validation_labels)}")
The two Linear layers register 42 trainable values through their weights and biases. Their matrix shapes are (8, 2) and (2, 8).
PyTorch accumulates gradients by default, so optimizer.zero_grad() clears the previous values before the next loss. loss.backward() computes parameter gradients and optimizer.step() applies them, as the PyTorch optimization tutorial explains.
Step 4: Evaluate without updating the weights
I kept the validation indices out of the optimizer loop. Those rows do not contribute gradients or weight updates, and model.eval() changes layers such as Dropout and BatchNorm while torch.inference_mode() skips gradient tracking.
Cross-entropy falls from 0.6426 to 0.0092, and predictions match all six training labels and both validation labels in this toy example.
I used matching validation indices for features and labels, so the 2/2 count checks target alignment as well as output shape. That two-row holdout still cannot estimate performance on a representative dataset.
Read shape, target, and device errors
A PyTorch forum question asks, “How could i change my output size as i expected?” Linear changes the final feature dimension, while the batch dimension passes through unchanged.
| Symptom | What to inspect | Correction |
|---|---|---|
| Linear reports incompatible matrix shapes | The last dimension of the input tensor | Set in_features to that dimension or reshape the feature data before the layer |
| CrossEntropyLoss rejects target values or shape | Logits should be (batch, classes). Class-index targets should be (batch,) with integer IDs from 0 through classes minus 1 | Use torch.long labels and keep one label for each input row |
| The model and input are on different devices | Check the device for the model, features, and labels | Move each to the same device before calling the model |
| Validation changes the weights | Check that validation batches are outside the optimizer loop | Use model.eval() and torch.inference_mode() for evaluation |
A new tensor or module is normally created on the CPU, not on a GPU. Selecting a device does not move objects that already exist, so move the model and all tensors before training when you choose an accelerator.
For a deeper loss breakdown, see this PyTorch loss-function reference. The complete PyTorch overview covers tensor operations that sit before the model.
A training score does not measure unseen data
Replace the hand-picked tensors with examples from your task and preserve the feature order. Keep representative validation rows out of optimizer updates, then run the same file.
python torch_nn_walkthrough.py
torch.nn questions that change the next step
torch.nn names a namespace, nn.Module is a base class, and nn.Linear creates a layer instance. Calling the model runs its forward pass, while calling the loss compares its scores with targets.
What is torch.nn in PyTorch?
torch.nn is PyTorch’s namespace for neural-network modules, layers, containers, activation functions, and losses. nn.Module is the base class for custom models.
Does torch.nn use the GPU by default?
No. New tensors and modules are normally created on the CPU. Move the model and its inputs to the same selected device to use an accelerator.
What does nn.Linear do?
nn.Linear applies an affine transformation to the last input dimension. Its in_features value must match that dimension, and its out_features value sets the output dimension.
Should I use nn.Module or torch.nn.functional?
Use an nn.Module for a parameterized layer that should register with the model. Use torch.nn.functional for an operation that does not need module-owned parameters.
Should I apply Softmax before CrossEntropyLoss?
No. CrossEntropyLoss expects unnormalized logits and class-index targets for this classification task, so pass the model scores directly.

