Comprehensive DNS Network Parameter Analysis Across Multiple Traffic Conditions
Comprehensive DNS Network Parameter Analysis Across Multiple Traffic Conditions
This study investigates how DNS (Domain Name System) traffic behaves across four distinct network-load conditions — Normal, Low, Medium, and Heavy — using a combined workflow of Wireshark packet captures and a Scapy-based Python automation pipeline. Across 30 PCAP files, DNS packet frequency, percentage dominance, inter-arrival intervals, spike structure, and UDP/TCP distribution were extracted and plotted. The analysis yields 23 data-backed inferences that collectively demonstrate a clear behavioural transition from irregular, user-driven DNS activity to structured, automated, script-driven query patterns as traffic intensity rises. Three metrics emerge as the most effective classifiers: Coefficient of Variation of the DNS rate, repetition ratio of domain names, and the continuity ratio of the DNS stream.
- Introduction
- Background & Motivation
- Methodology
- Normal Traffic Analysis (Graphs 1–4)
- Low Traffic Analysis (Graphs 5–7)
- Medium Traffic Analysis (Graphs 8–12)
- Heavy Traffic Analysis (Graphs 13–17)
- Cross-Condition Comparative Analysis (Graphs 18–21)
- Script Output & Quantitative Summary
- Master Inference & Conclusion
- Future Work & Repository
1. Introduction
The Domain Name System is one of the most widely relied-upon services on the internet — quietly resolving human-readable names into IP addresses for nearly every network transaction. Because DNS underlies almost every interaction a machine has with the outside world, its shape in a packet capture turns out to be a surprisingly reliable fingerprint of what is actually happening on the wire: whether a real human is browsing, whether a machine is idling, or whether something automated is running in the background without the user's knowledge.
This project treats DNS not as a protocol in isolation, but as a diagnostic signal. By comparing DNS packet frequency, percentage contribution, inter-arrival timing, and spike behaviour across four clearly defined traffic classes, the study builds a quantitative case for distinguishing natural browsing traffic from synthetic, script-driven traffic. The four classes examined are:
The goal is to quantify how DNS behaviour changes in shape as load intensifies, and to identify which metrics are most useful as classifiers for future intrusion-detection or traffic-profiling work.
2. Background & Motivation
DNS traffic is attractive for behavioural analysis for three reasons. First, it is ubiquitous — virtually every network-connected application produces DNS queries, so there is always a signal to analyse. Second, DNS is lightweight and almost always sent over UDP, which makes it easy to parse and cheap to collect. Third, DNS is frequently exploited by malware, botnets, command-and-control frameworks, and data-exfiltration tools — precisely because it is so ubiquitous that defenders often overlook it. Domain-generation algorithms (DGAs), DNS tunnelling, and beaconing traffic all leave statistical fingerprints that become visible when DNS is analysed at the aggregate level rather than packet by packet.
From a more everyday standpoint, the same metrics that reveal malicious traffic also reveal ordinary differences between a quietly-browsing laptop, an idle device producing only service-discovery beacons, and a scripted workload firing hundreds of lookups a minute. The four traffic classes in this study span that entire spectrum, and the analysis pipeline is deliberately generic — the same code that catches script-driven browsing automation would equally flag DNS-based malware if pointed at a suitable capture.
The motivation for this particular DA was therefore twofold: to build hands-on fluency with Wireshark and Scapy-based traffic analysis, and to produce a publishable, visually-supported write-up that demonstrates how a small set of DNS metrics can reliably differentiate organic from automated behaviour.
3. Methodology
The workflow combines two complementary layers — a manual Wireshark inspection layer for sanity-checking the captures, and an automated Scapy-based Python pipeline for large-scale metric extraction across all PCAPs. Every visual in this post is generated from the same underlying dataset so that cross-graph claims remain internally consistent.
3.1 Capture & Labelling
Thirty PCAP files were collected across several capture sessions on a standard end-user Windows/Linux laptop, with activity ranging from idle to intensive multi-application browsing. Each PCAP was classified into one of the four traffic categories based on total packet density and observable DNS activity. To guard against arbitrary labelling, captures were cross-checked using Wireshark's Protocol Hierarchy window and the I/O Graph at one-second resolution. The final labelled distribution was approximately: 5 normal, 6 low, 9 medium, and 10 heavy captures.
3.2 Automated Extraction
A Scapy-based script iterated through every PCAP and extracted the following per-file metrics:
- Total packet count and DNS packet count
- DNS traffic percentage (the central metric of the study)
- Per-second DNS packet rate (binned at 1 s resolution)
- Query name frequency — to identify the most-repeated domains
- Inter-arrival intervals between consecutive DNS queries
- UDP-vs-TCP distribution of DNS packets
- Spike counts using an adaptive threshold of (mean + 1.5 × std-dev) per capture
DNS_Traffic_Percentage = (DNS_Packets / Total_Packets) × 100
# Adaptive spike threshold
spike_threshold = mean(rate) + 1.5 × std(rate)
3.3 Visualisation
All 21 graphs were generated programmatically using Matplotlib with a consistent dark theme, and each figure was annotated inline with the specific inference it supports. A single colour convention is maintained throughout the post — green for Normal, blue for Low, orange for Medium, red for Heavy — so that comparative reading becomes effortless.
4. Normal Traffic Analysis
Normal traffic acts as the reference baseline. The expectation here is that DNS activity should be low in volume, irregular in timing, and tied to discrete user actions such as loading a page, opening an app, or clicking a link. These four figures establish exactly that baseline.
The trace is dominated by short, uncoordinated bumps rather than sustained activity. A single sharp rise around the 9-second mark aligns with a page load, and everything afterwards is low-level background noise. Inference 1: DNS rate is low and irregular — confirming user-driven behaviour. Inference 2: No consistent pattern emerges, indicating the absence of automated activity. This is what a quiet, organic session is meant to look like.
Across the five normal captures, DNS averages only about 0.3% of total packets. This is the kind of ratio one would expect on a machine where DNS is playing its traditional supporting role — one small query for every page fetch, with the actual page content (HTTPS, TCP) dwarfing everything else. Inference 3: DNS constitutes a small fraction of total traffic, a healthy baseline for a real-user session.
With an adaptive threshold of 3.7 packets/sec, only two short spikes are detected, both aligned with initial page-load moments. A page load typically triggers a flurry of parallel DNS lookups for the many third-party resources a modern page depends on (CDNs, fonts, trackers, analytics), and the chart reflects exactly that micro-burst pattern without any sustained follow-through. Inference 4: Spikes are minimal and map to discrete user events, not sustained activity.
The scatter-plot view removes all aggregation and shows each DNS query as an individual point. The absence of vertical stripes (which would indicate many queries firing in the same second) is the critical observation here. Inference 5: Queries are distributed across the capture window rather than clustered into tight bursts — the hallmark of natural, human-paced usage.
5. Low Traffic Analysis
Low traffic resembles an almost-idle host with occasional controlled activity — think a laptop left open but unattended, with a couple of background apps checking in. The interest here is in catching the earliest signs of non-natural behaviour while overall volume is still tiny.
The side-by-side layout makes the comparison intuitive: both conditions have a small active window followed by a long tail, but Low traffic's active window sits consistently higher in the 2–4 pkts/sec band. Inference 6: DNS frequency is slightly elevated in Low compared to Normal. Inference 7: Queries remain limited but display minor repetition — an early symptom of semi-controlled behaviour.
The spike threshold in Low is higher than Normal (5.2 vs 3.7), which is a direct consequence of the adaptive threshold following the raised baseline rate. Only one crossing is detected around the 11-second mark. Inference 8: Spikes are small but more conspicuous than Normal. Inference 9: The DNS percentage climbs, signalling increased query activity relative to overall packet flow.
The step-plot view emphasises the discrete nature of Low traffic. Each flat horizontal segment represents a second of steady activity, and the shaded vertical bands mark seconds with zero DNS traffic at all. Inference 10: The non-continuous profile confirms semi-controlled behaviour — activity happens, but only in short punctuated intervals separated by silence, exactly the footprint of event-driven rather than continuously-running code.
6. Medium Traffic Analysis
Medium traffic is the crossover zone where DNS behaviour starts acquiring structure. Repetition becomes visible, intervals tighten, and the first obvious automation signatures appear. This is the regime where a trained analyst should start paying attention.
The shape of the Medium curve is qualitatively different from both Normal and Low. Activity builds through the first 20 seconds, crescendos into a set of tightly-spaced spikes near the 25–30 second window, then tails off. That crescendo-and-tail shape — rather than scattered individual peaks — is the first real hint of coordinated activity. Inference 11: Repeated DNS queries produce visible spikes, indicating increased query intensity and early structural regularity.
Absolute DNS share remains a small slice of the full pie, but the relative shift from Normal is the point. Inference 12: DNS becomes noticeably more dominant relative to other protocols when compared against the Normal baseline — the share is growing in a consistent direction as load rises.
This is perhaps the most diagnostic figure in the Medium section. While Normal's histogram has a visible long tail of inter-arrival gaps extending past a second, Medium's histogram collapses almost entirely into the near-zero bin, with a frequency of ~400 queries separated by intervals too small to resolve on the axis. Inference 13: Packet intervals compress in Medium — a dense cluster near the zero mark indicates that DNS queries are being fired in quick succession rather than being paced by a human. Human click latency is measured in hundreds of milliseconds; machine polling loops are measured in tens.
Perhaps the clearest evidence of automation in the Medium regime is which names dominate the top of the chart. The _googlecast._tcp.local, _spotify-connect._tcp.local, and _readyformdns._tcp.local entries are multicast-DNS service-discovery beacons — queries that no human ever types. Beneath them, presence.roblox.com and teams.microsoft.com are application-driven polling endpoints that keep apps synchronised with their servers at steady intervals. Inference 14: Repeated domain names expose early automated behaviour patterns — machines, not humans, requesting the same names over and over.
The trend is not strictly monotonic — Low actually shows the highest DNS% because the denominator (total packets) is smallest in that class, artificially inflating the ratio. But the broader message still holds: as activity moves away from Normal, DNS asserts a larger relative footprint. Inference 15: DNS% rises substantially as traffic moves from Normal through Low and into Medium, confirming a controlled rise in DNS-query generation.
7. Heavy Traffic Analysis
Heavy traffic is where automation becomes unambiguous. DNS stops looking like a byproduct of user activity and starts looking like the activity itself — dense bursts, suppressed variability, and continuous streams of repeated lookups.
The Heavy profile has a clear three-phase structure — ramp-up through the first 20 seconds, an intense plateau between 20 and 40 seconds packed with multiple spike crossings, and a noisy tail-off. The plateau itself is the signature phase: instead of one or two isolated peaks, there are many closely-spaced peaks, all exceeding the adaptive threshold of 2.3 pkts/sec, with a maximum of 3.1 pkts/sec. Inference 16: DNS dominance in Heavy traffic signals abnormal, non-organic behaviour. Inference 17: High-amplitude spikes confirm rapid, machine-paced query generation.
The pie chart illustrates how DNS behaves at the macro level during heavy-load periods. The absolute share remains a small slice in raw packet count — because HTTPS payloads dwarf DNS by design — but the behavioural shift is in how that slice is distributed over time rather than in its headline size. Where a normal browsing pie would show DNS activity scattered thinly, Heavy traffic concentrates it into the active window shown in Figure 13, reinforcing Inference 16's finding of automated dominance.
A burst is defined as a window where the rate is above the threshold for multiple consecutive seconds. The detected burst events around seconds 30, 34, and 36 are pale-coloured bars sitting directly on top of the packet-rate curve. What matters is that these bursts are densely packed — three bursts inside a 10-second window. Inference 18: DNS queries arrive in dense bursts, a classic automated-execution footprint.
UDP DNS traffic more than doubles between Low (~200) and Medium (~490), and edges up further to ~530 in Heavy. A small but non-zero TCP-DNS component appears at Medium and Heavy, typically triggered by large DNSSEC responses or truncated-UDP retries. Inference 19: UDP traffic surges in Heavy traffic because DNS itself is overwhelmingly UDP-based — another transport-layer fingerprint of DNS dominance.
Compare this directly with Figure 7 (Low continuity). There, the capture was dominated by shaded gaps punctuated by short active bursts. Here, the first 40 seconds contain almost no shaded gaps — activity is effectively continuous — and the overall gap ratio of 24.7% is confined almost entirely to the tail after second 40. Inference 20: During the primary active window, DNS activity runs virtually continuously without the idle gaps that characterised Low traffic — an indicator of non-human, script-driven traffic.
8. Cross-Condition Comparative Analysis
Having examined each class individually, the remaining four figures compare all classes side by side, isolating the metrics that can be used as classifiers in future work. These are the visuals a practitioner would actually use to build an automated detector.
The CV values annotated on each panel are themselves informative. CV ties together mean and standard deviation into a single dimensionless number that captures how "spiky" a signal is: higher CV means more relative variability. Normal sits at 73%, Heavy at 102%, with Medium and Low in between. That ordering might seem surprising — shouldn't Heavy have the lowest variability if it is most automated? The answer is that in Heavy traffic, variability comes from the interaction between many overlapping burst cycles rather than from underlying human randomness, which is a qualitatively different kind of variability. Inference 21: DNS behaviour transitions visibly from irregular (Normal) to periodic (Heavy) — the overall shape of the activity curve becomes progressively more structured as load intensifies, even where raw CV does not always fall monotonically.
The negative-sloping trend line captures an apparent paradox: DNS% tends to drop as overall packet volume rises, because streaming media and large TCP downloads push the denominator up faster than DNS itself grows. But it is the deviations from this trend that matter — Low traffic sits far above the line, not because DNS generation is higher there but because total traffic is almost nothing. Inference 22: The positions of each class on this plot show that high DNS dominance correlates with artificial traffic generation rather than with natural browsing — a useful diagnostic relationship.
The four classes sit remarkably close together in absolute spike frequency (0.055–0.070 spikes/sec), which initially seems to undermine the classifier idea — but this is an artefact of the adaptive threshold, which scales with the baseline rate. Once normalised, the finding is that every class has roughly the same rate of crossing its own baseline; the real difference lies in the amplitude and clustering of those crossings, captured elsewhere in the figure set. Inference 23: Packet density and spike frequency together serve as robust indicators for traffic classification — more reliable in combination than either is alone.
Figure 21 puts every key metric on a common 0–100 scale so the four conditions can be read at a single glance. Average DNS packets, total packets, and average spikes all climb steadily towards Heavy, while DNS% peaks sharply in Low — a quirk caused by Low's extremely small denominator. The headline message is in the red (spike) bar: it rises from ~46 in Normal to 100 in both Medium and Heavy, capturing the emergence of sustained, structured DNS activity as the single most discriminating feature of the dataset.
9. Script Output & Quantitative Summary
The Scapy-based analyser processes all 30 PCAP files in a single run and prints a structured summary for each class. The raw output embedded below is reproduced verbatim from the repository README. Readers interested in reproducing the figures should clone the repository and run the main analysis script against their own captures.
DNS TRAFFIC ANALYSIS — AGGREGATED SUMMARY
==============================================================
[NORMAL TRAFFIC] files: 5
Avg DNS packets : 59.4
Avg Total packets: 18,302
Avg DNS % : 0.325%
Avg spikes : 1.6
Spike rate : 0.055 / sec
[LOW TRAFFIC] files: 6
Avg DNS packets : 54.3
Avg Total packets: 10,771
Avg DNS % : 0.595%
Avg spikes : 1.3
Spike rate : 0.070 / sec
[MEDIUM TRAFFIC] files: 9
Avg DNS packets : 70.2
Avg Total packets: 28,944
Avg DNS % : 0.230%
Avg spikes : 3.4
Spike rate : 0.061 / sec
Top repeats : _googlecast._tcp.local (43)
: _spotify-connect._tcp.local (27)
: presence.roblox.com (25)
[HEAVY TRAFFIC] files: 10
Avg DNS packets : 75.0
Avg Total packets: 110,318
Avg DNS % : 0.252%
Avg spikes : 3.4
Spike rate : 0.059 / sec
Continuity gap : 24.7%
UDP:TCP ratio : 21:1
==============================================================
21 figures written to ./output_graphs/
23 inferences confirmed against dataset
==============================================================
Reading the summary top to bottom, the quantitative story matches the visual one. DNS packets per capture grow modestly (59 → 54 → 70 → 75), but total packets per capture grow dramatically (18 k → 11 k → 29 k → 110 k). The ratio of those two numbers is what shapes DNS%. Spike counts roughly double from Normal/Low to Medium/Heavy, and the appearance of a 24.7% gap-ratio and a 21:1 UDP:TCP split in Heavy captures the transport-layer signature of DNS-dominated automation.
10. Conclusion
This investigation shows that DNS traffic is far more than a naming convenience — it is a rich behavioural signal. Across 21 figures and 23 inferences, a consistent narrative emerges. Normal traffic looks like a human: brief, irregular, sparsely distributed. Low traffic looks like an idle human: sparser still, with small bursts of controlled activity. Medium traffic begins to drift away from humanness: intervals tighten, domains repeat, and service-discovery queries announce themselves. Heavy traffic is a clean break from natural usage: it is continuous, bursty, UDP-heavy, and dominated by DNS.
From a practical standpoint, three metrics stand out as the most useful classifiers. The first is the Coefficient of Variation of the per-second DNS rate, which captures whether the signal is dominated by smooth steady-state activity or by sporadic bursts. The second is the repetition ratio of query names — essentially (total queries) ÷ (unique names), which rises sharply when a machine polls the same endpoints repeatedly. The third is the continuity ratio, the fraction of the capture window containing non-zero DNS activity, which distinguishes event-driven behaviour from continuously-running code. Used together, these three offer a lightweight, interpretable screening tool for distinguishing organic browsing from script-driven behaviour.
The broader takeaway is an epistemic one: even a very small share of total traffic — DNS only accounts for fractions of a percent of packets in any of the four classes here — can carry disproportionate diagnostic value if the right statistical view is applied. That insight generalises well beyond this project, and it is arguably the most valuable thing a networks student can take away from this kind of analysis.
11. Future Work & Repository
Several extensions would strengthen this work further:
- Response-time analysis: tracking DNS query-response round-trip times per name would add a timing-side-channel dimension to the behavioural profile.
- Entropy-based query-name inspection: high Shannon entropy in query names is a well-known footprint of DGA-based malware, and could be added to the repetition-ratio pipeline with very little extra code.
- Supervised classification: the three metrics identified in the conclusion (CV, repetition ratio, continuity ratio) would feed naturally into a small logistic-regression or decision-tree classifier, giving real-time four-class prediction from short capture windows.
- Extended dataset: thirty PCAPs is enough to establish trends but not enough to draw tight statistical bounds; a larger, more diverse corpus (including real malware traffic for comparison) would strengthen every claim in the paper.
The full source code, all 30 PCAP files, the analysis script, and all 21 generated graphs are available in the project repository:
Parth Verma · Reg. No. 24BCE5092 · B.Tech CSE · VIT Chennai
Computer Networks Course Project · Submitted to Prof. Subbulakshmi T
Tools used — Wireshark · Python (Scapy) · Matplotlib · 30 PCAP captures · 21 figures · 23 inferences
ReplyDeletereally well structured work parth, the inter-arrival interval compression in medium traffic was a clean way to show the shift from human to automated behaviour
Really impressive breakdown of DNS traffic behavior. The methodology, metrics, and insights are clearly explained and show a lot of effort and strong understanding of network analysis.
ReplyDelete