X1 - Midterm Alternative: Anatomy of a Page Load
Week 8 · 160 points · take home · individual · about 10 hours · submit in Canvas
Overview
This assignment is an alternative to the midterm exam. It covers the same material, chapters 1 through 4, but instead of answering questions about delay, DNS, TCP and NAT from memory you are going to measure every one of them on real networks and explain what you see.
You will use a command line AI agent to build a small C program called pageload that fetches a URL over plain HTTP and times each step along the way: the DNS lookup, the TCP handshake, the wait for the first byte, and the transfer. Then you run it from three different networks against servers in Utah, California, New York, Germany and Japan, predict what the numbers should be before you look, and explain the difference between your predictions and reality using the textbook.
The agent writes the code. You do the networking. Almost all of the points are in the analysis, because that is the part an exam would test.
How this counts
X1 and the Midterm Exam are both in the Midterm assignment group (25% of your final grade), and that group drops the lowest score. So you have three options:
- Take the midterm exam only
- Do X1 only
- Do both, and the better of the two scores counts
X1 is due at the same time the midterm closes, Friday of week 8 at 11:59pm, and there is no grace period. It replaces an exam, so the 2 day grace period for homework in the syllabus does not apply. Work submitted 1 second late is treated the same as work submitted 1 day late.
This is individual work
X1 replaces an exam, so it is individual. Do not share code, measurements, or any part of your report with anyone else in the class. You can talk about the textbook with classmates, but your numbers and your explanations must be your own. Two reports with the same measurements are both a zero.
Learning Outcomes
- 1.2 Compare circuit switching and packet switching, and account for delay, loss, and throughput in a packet-switched network
- 2.3 Describe how DNS resolves a name, and why it is structured as a distributed hierarchical database
- 3.2 Analyze TCP connection management, flow control, and congestion control
- 4.1 Describe the data plane: forwarding, the IP datagram, addressing, and NAT
- 4.4 Subnet an address block with CIDR and compute the resulting ranges
- 7.3 Compile and run code on at least two different systems
AI First
This assignment is AI first. That means you start every task by asking your agent, not by opening a blank file. The course AI policy has no restrictions, and this assignment leans into that on purpose: the skill being tested is whether you can drive an agent to a correct result and then tell whether what it gave you is actually correct.
- Use a command line agent. GitHub Copilot CLI (free with GitHub Education, see M1), Claude Code, OpenAI Codex CLI, Gemini CLI, or similar. It has to be a tool that runs in your terminal, edits files, and runs commands, not a chat window in a browser.
- Save every session. Export each session to Markdown before you close it (
/share file ./session-1.mdin Copilot CLI,/exportin Claude Code). If you forget,copilot --resumelets you reopen it. You will submit all of them. - You can use the agent for the analysis too. But every answer must be built on your own measurements, and you are responsible for every claim in your report. An answer that does not use your numbers earns no points for that question, no matter how well written it is.
- Do not pay for anything. Copilot CLI through GitHub Education is all you need. If it says you are out of requests, stop and email me.
Task 1 - Setup
You need three vantage points, three different networks to measure from:
- Your laptop, on the network you normally use (home, apartment, campus Wi-Fi, whatever it is). A Mac is fine as it is, and on Windows use WSL.
- Onyx, which is on campus in Boise.
- A GitHub Codespace, which runs in a Microsoft data center somewhere.
Do the agent work on your laptop. Make a directory for the assignment, make it a private GitHub repository, and start your agent inside it:
mkdir cs425-x1
cd cs425-x1
git init
copilotOnce pageload builds, push the repository to GitHub and open a Codespace on it. Onyx you reach with ssh onyx, the key and the one-word shortcut you set up in A2. The agent does not need to run on Onyx or in the Codespace, only your program does.
This is what A2 was for
An agent cannot type your password. Because ssh onyx logs in with your key and no prompt, your agent can copy your code to Onyx, build it, run your measurements and copy the results back, all from your laptop, while you watch and approve each command:
ssh onyx mkdir -p cs425-x1
scp -r pageload.c Makefile measure.sh .git onyx:cs425-x1/
ssh onyx 'cd cs425-x1 && make && ./measure.sh onyx'
scp onyx:cs425-x1/measurements-onyx.txt .The .git directory goes along so git rev-parse in measure.sh can print the commit you measured. The connection multiplexing from A2 step 4 means only the first of those pays for a full login, and the rest reuse the open connection. If ssh onyx still asks for a password, go back and finish A2 steps 1 through 3 before you start anything else. It takes about 15 minutes and you will save that many times over this week.
Task 2 - Build pageload with your agent
Give your agent the specification below. You can paste it in, break it into smaller prompts, or describe it in your own words. Read each command the agent asks to run before you approve it, the same way you did in M1.
pageload [-4 | -6] [-n trials] [-o file] http://host[:port][/path]- Accept only
http://URLs. Rejecthttps://and anything malformed with a useful error and a nonzero exit status. The port defaults to 80 and the path to/. -4uses only IPv4 and-6only IPv6. With neither, use whatevergetaddrinforeturns first.-nrepeats the whole fetch that many times (default 1), including the DNS lookup.-owrites the response body of the last trial to a file.- For each trial, do the following with the sockets API (no
curl, no libcurl, no HTTP library):- Resolve the host with
getaddrinfo, and time it. - Open a TCP socket and
connect, and time only theconnectcall. - Send an HTTP/1.1
GETwith aHostheader andConnection: close. - Time from sending the request to the first byte of the response.
- Read until the server closes the connection, and time that.
- Resolve the host with
- For each trial, print: the DNS time and the address it resolved to, the connect time and the full 4-tuple (local address and port, remote address and port, using
getsockname), the first byte time, the transfer time with the number of body bytes and the throughput in Mbit/s, and the HTTP status line. - After the last trial, print the minimum and median connect time.
- Time everything with
clock_gettime(CLOCK_MONOTONIC, ...). - A
Makefilebuilds it with-Wall -Wextra, with no warnings on both Codespaces and Onyx, and has acleantarget.
The exact layout of the output is up to you. Here is what mine looks like, run from my laptop on a home network:
$ ./pageload -4 -n 3 http://ftp.jaist.ac.jp/debian/README
trial 1 http://ftp.jaist.ac.jp/debian/README
dns 15.7 ms ftp.jaist.ac.jp -> 150.65.7.130:80 (IPv4)
connect 260.1 ms 192.168.4.42:51549 -> 150.65.7.130:80
first byte 286.4 ms
transfer 0.7 ms 1198 bytes, 13.93 Mbit/s
status HTTP/1.1 200 OK
trial 2 http://ftp.jaist.ac.jp/debian/README
dns 1.1 ms ftp.jaist.ac.jp -> 150.65.7.130:80 (IPv4)
connect 247.2 ms 192.168.4.42:51552 -> 150.65.7.130:80
first byte 280.2 ms
transfer 0.0 ms 1198 bytes, 273.83 Mbit/s
status HTTP/1.1 200 OK
trial 3 http://ftp.jaist.ac.jp/debian/README
dns 3.6 ms ftp.jaist.ac.jp -> 150.65.7.130:80 (IPv4)
connect 253.3 ms 192.168.4.42:51556 -> 150.65.7.130:80
first byte 313.5 ms
transfer 0.0 ms 1198 bytes, 199.67 Mbit/s
status HTTP/1.1 200 OK
connect over 3 trials: min 247.2 ms, median 253.3 msAnd a bigger download, where the transfer time finally matters:
$ ./pageload -4 http://mirrors.xmission.com/debian/dists/trixie/main/binary-amd64/Packages.gz
trial 1 http://mirrors.xmission.com/debian/dists/trixie/main/binary-amd64/Packages.gz
dns 7.9 ms mirrors.xmission.com -> 198.60.22.13:80 (IPv4)
connect 31.1 ms 192.168.4.42:51561 -> 198.60.22.13:80
first byte 34.0 ms
transfer 545.8 ms 13332733 bytes, 195.42 Mbit/s
status HTTP/1.1 200 OK
connect over 1 trials: min 31.1 ms, median 31.1 msTest it before you trust it. Compare the address it prints with dig, and the body it saves with -o against the same URL in a browser. On Onyx you also still have the probe function from A2, which times the same connect with curl, so ssh onyx probe mirrors.xmission.com 80 and pageload on Onyx should report connect times in the same ballpark. If something is wrong, tell the agent what happened and let it fix it, and leave that back and forth in the log.
Task 3 - Predict
Before you run a single measurement, write your predictions down in report.md. Then paste them into your agent session and ask it to comment on them. That puts your predictions in the session log, timestamped, before any of the measurements. A prediction written after you have seen the numbers is not a prediction.
Predict the following. A wrong prediction costs you nothing, a missing one does.
- The minimum connect time from your laptop to each of the five mirrors in the table below, computed from distance alone: the great circle distance there and back at the speed of light in fiber, 2 x 10^8 m/s (200 km per millisecond).
- Which mirror will have the longest connect time from the Codespace, and why.
- Whether the DNS time changes between trial 1 and trial 2, and why.
- Which of your three vantage points are behind a NAT.
- Whether downloading the same 13 MB file from Utah and from Japan will take about the same time, and why.
Task 4 - Measure
Run every target below from all three vantage points, with -4 and -n 5:
| Target | URL | Location |
|---|---|---|
| CDN | http://example.com/ | served by a content delivery network |
| Utah | http://mirrors.xmission.com/debian/README | Salt Lake City, UT |
| California | http://mirrors.ocf.berkeley.edu/debian/README | Berkeley, CA |
| New York | http://mirrors.rit.edu/debian/README | Rochester, NY |
| Germany | http://ftp.fau.de/debian/README | Erlangen, Germany |
| Japan | http://ftp.jaist.ac.jp/debian/README | Nomi, Japan |
Then, from all three vantage points, download the same file from Utah and from Japan with -4 -n 3:
/debian/dists/trixie/main/binary-amd64/Packages.gzIt is about 13 MB, and from far away it can take 30 seconds or more per trial, so be patient. Finally, from each vantage point, ask a server what your public address is:
./pageload -4 -o public-ip.txt http://checkip.amazonaws.com/
cat public-ip.txtYou are not going to type all of that three times. Have your agent write a script called measure.sh that does it for you, to the specification below.
What measure.sh must do
- Take the vantage point as its only argument.
./measure.sh laptop,./measure.sh onyxor./measure.sh codespace. Anything else, or no argument, prints a usage message and exits with status 1. - Build first. Run
makebefore anything else and stop if the build fails, so a run never measures with a stale binary. - Keep every target in one list at the top of the script, so you (and I) can see what it measures without reading the whole thing.
- Write a header before the first measurement: the vantage point name, the date and time in UTC (
date -u),hostname,uname -a, the short git commit of the code being run (git rev-parse --short HEAD), and the output ofip -4 addr. You need that last one for question 15. If your laptop does not haveip, have your agent find the command that shows the same thing. - Run these sections, in this order, each one starting with a marker line such as
=== latency utah ===and the time it started:- Public address. Fetch
http://checkip.amazonaws.com/with-oand print the address it returned (question 13). - Latency. Every target in the table above, with
-4 -n 5. - Downloads.
Packages.gzfrom Utah and from Japan, with-4 -n 3. - Two at once. Two copies of
pageloadagainst the Utah mirror at the same time, each writing to its own temporary file, then both files printed one after the other so the output does not interleave (question 11). - DNS trace.
dig +trace ftp.jaist.ac.jp(question 7). It fails on Onyx, which blocks DNS to every server but the campus one (you saw that in A1), and ifdigis not installed it cannot run at all. Either way, print a note saying so and keep going.
- Public address. Fetch
- Keep going when a target fails. A mirror that is down or slow should not end the run. Print the exit status of the failed
pageloadand move on to the next target. - Save everything. Both stdout and stderr go to the screen and to
measurements-<vantage>.txt(usetee), replacing the file from any earlier run. - Finish with a summary line: how long the whole run took (bash keeps a count in
$SECONDS) and how many targets failed. - Nothing else. No
sudo, no installing packages, and no editing of the results. It is a plainbashscript that starts with#!/usr/bin/env bash.
The finished file has a shape like this (the ... is where the pageload output goes):
=== x1 measure.sh vantage=onyx ===
date 2026-10-14T18:02:11Z
hostname onyx
uname Linux onyx 5.14.0 ... x86_64 GNU/Linux
commit 3f9c2a1
...output of ip -4 addr...
=== public address 18:02:12 ===
...
=== latency cdn 18:02:13 ===
...
=== latency utah 18:02:15 ===
...
=== download japan 18:04:40 ===
...
=== two at once 18:06:02 ===
...
=== dns trace 18:06:04 ===
...
=== done in 241 s, 0 failed ===A full run takes about 5 to 10 minutes, most of it the Japan downloads. Run the whole thing at least once per vantage point and submit the raw output. Do not edit it: the commit and the timestamps in the header are how I match it to your code and your session log.
You also need to know roughly where each vantage point is to work out the distances. For your laptop and Onyx that is easy. For the Codespace, your region setting at github.com/settings/codespaces is a good start. Say how you decided.
Task 5 - Analyze
Answer each question in report.md, under a heading with its number. Show your work for every calculation, cite the textbook section that backs up each explanation, and use your measurements. A few sentences and a table is the right size for most questions.
Chapter 1 - Delay and throughput
- Distance versus delay. Make a table with a row for every mirror and vantage point: the great circle distance, the propagation round trip time at 2 x 10^8 m/s, your minimum connect time, and the ratio of the two.
- Where the extra time goes. Every ratio is above 1. Explain why using the four components of nodal delay from section 1.4, plus the fact that packets do not travel in a straight line. Which component dominates for Japan, and which for Utah?
- Minimum, not mean. Why does question 1 use the minimum connect time instead of the mean? Use the spread of your own 5 trials to back it up.
- Transmission delay. A SYN segment is about 64 bytes on the wire. Compute its transmission delay on a 10 Mbit/s link and on a 1 Gbit/s link, and compare both to your propagation delay to Utah. What does that tell you about which delay matters for small messages?
- Throughput. For the
Packages.gzdownloads, compare the throughput from Utah and from Japan at each vantage point. Is your access link the bottleneck in both cases? How do you know?
Chapter 2 - The application layer
- DNS caching. Compare the DNS time of trial 1 with trials 2 through 5 for one mirror at each vantage point. Explain the difference in terms of the local DNS server and caching.
- The DNS hierarchy. Using the DNS trace from a vantage point where it worked, list, in order, the servers that were asked and what each one answered (a referral or the address). Map each one to its place in the hierarchy from section 2.4.
- The CDN. Did DNS give
example.comthe same address at all three vantage points? How does its connect time compare with the mirrors? Use section 2.6 to explain how the CDN got a server close to you, and say which technique your evidence points to. - Counting round trips. The first byte time is roughly one RTT plus the time the server spends answering. Estimate the server time for two mirrors. Then count the round trips
pageloadspends from the start of the DNS lookup to the last byte of the README, and compare that with the non-persistent HTTP estimate in section 2.2. You did the same thing for a cold SSH login in A2 step 4, so use that as your model.
Chapter 3 - The transport layer
- The handshake. Your connect time is about one RTT, not one and a half. Which segments have been sent and received when
connectreturns? Why does the client not have to wait for the third one? - Demultiplexing. Use the "two at once" section of one of your measurement files, where two copies of
pageloadran against the same mirror at the same time. Write down both 4-tuples. What is different between them, who picked it, and how does the server tell the two connections apart when both arrive at port 80? (section 3.2) - Slow start. For
Packages.gz, assume an MSS of 1,460 bytes, an initial congestion window of 10 segments, no loss, and a window that doubles every RTT. How many RTTs does it take to send the whole file? Using your minimum connect time as the RTT, what is the shortest possible transfer time from Utah and from Japan? Compare with what you measured. Which path comes closer to the estimate, and what is holding the other one back? (sections 3.5 and 3.7)
Chapter 4 - The network layer
- Private and public addresses. For each vantage point, give the local address from the 4-tuple and the public address from
checkip. Which local addresses are private (RFC 1918)? Which vantage points are behind a NAT, and what is your evidence? - The NAT table. For one vantage point that is behind a NAT, write the NAT translation table entry for one of your connections, both the WAN side and the LAN side, the way section 4.3 draws it. Fill in everything you can observe, and for anything you cannot, say why you cannot see it from where you are.
- Subnetting. For each vantage point, find the interface address and prefix length in the header of its measurement file. Compute the network address, the broadcast address, the usable host range, and the number of usable hosts. Show the binary for one of them.
Task 6 - Catch the agent
Find one place where your agent was wrong, incomplete, or made a suggestion you did not take. Quote it from the session log, say how you checked it (a man page, the textbook, a measurement), and give the correct answer.
If the agent never got anything wrong, pick the most important claim it made about your measurements, verify it independently, and say how. "It was always right" with nothing behind it earns no points.
What to submit
Upload these files directly to the Canvas assignment. Do not zip them.
| File | What it shows |
|---|---|
pageload.c and Makefile | The program your agent helped you build |
measure.sh | The script that ran your measurements |
measurements-laptop.txt, measurements-onyx.txt, measurements-codespace.txt | The raw output, unedited |
session-*.md | Every exported agent session |
report.md | Your predictions, your answers to Task 5, and Task 6 |
Rubric
| # | Criterion | Points |
|---|---|---|
| 1 | pageload meets the specification and builds with no warnings on Codespaces and Onyx | 20 |
| 2 | Session logs show the agent built the tool and the predictions were logged before measuring | 10 |
| 3 | Measurements are complete: three vantage points, every target, every download, raw output | 15 |
| 4 | Chapter 1: distance table, delay components, minimum versus mean, transmission delay | 25 |
| 5 | Chapters 1 and 3: throughput, the bottleneck, and the slow start estimate | 20 |
| 6 | Chapter 2: DNS caching, the DNS hierarchy, the CDN, and counting round trips | 25 |
| 7 | Chapter 3: the handshake and demultiplexing | 15 |
| 8 | Chapter 4: private and public addresses, the NAT table, and subnetting | 20 |
| 9 | Catch the agent | 10 |