Finding Repeated Log Entries in Python
Find log entries with the same time, source, destination and port. The report keeps their line numbers so you can inspect the repeats without deleting anything.
You'll need Python 3, a plain-text editor and a command window. No extra packages are required.
1. Create the log file
Use a tab-separated text file with four columns: time, source address, destination address and destination port. The header must be Time, Source, Destination, Port, with actual tabs between the names. A .tsv file is a plain-text table that uses tabs to separate columns.
Write times as YYYY-MM-DDThh:mm:ssZ. The Z means Coordinated Universal Time (UTC), the shared reference time rather than a local clock. This reader accepts whole seconds in that exact form, not fractional seconds, local offsets such as +01:00 or leap-second values. Addresses use Internet Protocol version 4 (IPv4), four dot-separated numbers such as 192.0.2.10, without extra leading zeros. Ports are whole numbers from 1 to 65535.
The sample addresses are reserved for documentation in RFC 5737, a published document describing addresses to use in examples. This is the reader's own input format, not a direct export from a network-monitoring product.
Save this as MakeRepeated.py in a practice folder:
from pathlib import Path
Rows = [
"Time\tSource\tDestination\tPort\n",
"2026-09-30T09:00:00Z\t192.0.2.10\t198.51.100.20\t443\n",
"2026-09-30T09:00:00Z\t192.0.2.10\t198.51.100.20\t0443\n",
"2026-09-30T09:00:01Z\t192.0.2.10\t198.51.100.20\t443\n",
"2026-09-30T09:00:00Z\t192.0.2.10\t198.51.100.20\t70000\n",
]
Path("Repeated.tsv").write_bytes("".join(Rows).encode("ascii"))
Run:
python3 MakeRepeated.py
Use your installation's Python 3 command if it is not named python3. This creates or replaces Repeated.tsv, so use a practice folder without a file you need to keep. The \t sequences write actual tabs and \n writes line feeds; you do not need to insert tabs in your editor.
Lines 2 and 3 share parsed values, the values extracted after format checking. Ports 443 and 0443 become the same whole number. Line 4 has a different time; line 5 has an unsupported port.
2. Add the reader
The complete reader is included below so you can run this guide on its own.
Save this as LogReader.py. This reusable source file checks each row and keeps its physical line number, either with accepted values or a reason it could not be read.
from dataclasses import dataclass as DataClass
from datetime import datetime as DateTime
from pathlib import Path
@DataClass
class LogEntry:
Number: int
Original: str
EventTime: str = ""
Source: str = ""
Destination: str = ""
Port: int = 0
Problem: str = ""
Accepted: bool = False
def Digits(Value, Maximum):
if not Value:
raise ValueError("empty number")
if any(Character not in "0123456789" for Character in Value):
raise ValueError("number needs ordinary digits")
Number = int(Value)
if Number > Maximum:
raise ValueError("number outside supported range")
return Number
def CheckAddress(Value):
Parts = Value.split(".")
if len(Parts) != 4:
raise ValueError("expected four IPv4 address parts")
for Part in Parts:
if len(Part) > 3 or (len(Part) > 1 and Part.startswith("0")):
raise ValueError("use canonical dotted IPv4 addresses")
Digits(Part, 255)
def CheckTime(Value):
if len(Value) != 20:
raise ValueError("expected time YYYY-MM-DDThh:mm:ssZ")
if any(Value[Index] != Mark for Index, Mark in
((4, "-"), (7, "-"), (10, "T"), (13, ":"), (16, ":"), (19, "Z"))):
raise ValueError("expected time YYYY-MM-DDThh:mm:ssZ")
NumberText = Value[0:4] + Value[5:7] + Value[8:10] + Value[11:13] + Value[14:16] + Value[17:19]
if any(Character not in "0123456789" for Character in NumberText):
raise ValueError("time needs ordinary digits")
Year = Digits(Value[0:4], 9999)
Month = Digits(Value[5:7], 12)
Day = Digits(Value[8:10], 31)
Hour = Digits(Value[11:13], 23)
Minute = Digits(Value[14:16], 59)
Second = Digits(Value[17:19], 59)
try:
DateTime(Year, Month, Day)
except ValueError:
raise ValueError("date does not exist") from None
try:
DateTime(Year, Month, Day, Hour, Minute, Second)
except ValueError:
raise ValueError("time does not exist") from None
def ParseEntry(Entry):
try:
if len(Entry.Original) > 1024:
raise ValueError("line exceeds 1024 bytes")
if any(Character != "\t" and not 32 <= ord(Character) <= 126
for Character in Entry.Original):
raise ValueError("unsupported byte in practice format")
Fields = Entry.Original.split("\t")
if len(Fields) != 4:
raise ValueError("expected exactly four tab-separated fields")
CheckTime(Fields[0])
CheckAddress(Fields[1])
CheckAddress(Fields[2])
Entry.Port = Digits(Fields[3], 65535)
if Entry.Port == 0:
raise ValueError("port must be 1 to 65535 in this format")
Entry.EventTime, Entry.Source, Entry.Destination = Fields[:3]
Entry.Accepted = True
except ValueError as Error:
Entry.Problem = str(Error)
return Entry
def ReadLog(FileName):
with Path(FileName).open("rb") as Input:
Data = Input.read(1048577)
if len(Data) > 1048576:
raise ValueError("input exceeds 1 MiB teaching limit")
if not Data:
raise ValueError("empty input")
Lines = Data.split(b"\n")
if Lines[-1] == b"":
Lines.pop()
Lines = [Line[:-1] if Line.endswith(b"\r") else Line for Line in Lines]
if Lines[0] != b"Time\tSource\tDestination\tPort":
raise ValueError("unsupported header")
if len(Lines) - 1 > 1000:
raise ValueError("more than 1000 data lines")
return [ParseEntry(LogEntry(Number, Line.decode("latin-1")))
for Number, Line in enumerate(Lines[1:], start=2)]
The reader handles up to 1,000 data lines in a file of at most 1 MiB, roughly one million bytes, and rejects data lines longer than 1,024 bytes. Its fields use ordinary printable English-character bytes and tabs. Use a saved file that will not change while it is read.
3. Add the repeat reviewer
Save this as FindRepeatedEntries.py:
import json as Json
import sys as Sys
from LogReader import ReadLog
def FindRepeatedEntries(FilePath):
Groups = {}
Rejected = []
Accepted = 0
for Entry in ReadLog(FilePath):
if not Entry.Accepted:
Rejected.append({"Line": Entry.Number, "Reason": Entry.Problem})
continue
Accepted += 1
Key = (Entry.EventTime, Entry.Source, Entry.Destination, Entry.Port)
Groups.setdefault(Key, []).append(Entry.Number)
Repeated = []
for Key in sorted(Groups):
if len(Groups[Key]) > 1:
Repeated.append({"Time": Key[0], "Source": Key[1], "Destination": Key[2],
"Port": Key[3], "Lines": Groups[Key]})
return {"AcceptedEntries": Accepted, "RepeatedGroups": Repeated, "RejectedLines": Rejected}
def Main():
if len(Sys.argv) != 2:
print("Usage: python3 FindRepeatedEntries.py input.tsv", file=Sys.stderr)
return 2
try:
Report = FindRepeatedEntries(Sys.argv[1])
except (OSError, ValueError) as Problem:
print(f"Review stopped: {Problem}", file=Sys.stderr)
return 2
print(Json.dumps(Report, ensure_ascii=True, indent=2))
return 1 if Report["RepeatedGroups"] or Report["RejectedLines"] else 0
if __name__ == "__main__":
Sys.exit(Main())
Groups is a dictionary, a collection for looking up values by a key. The key combines time, source, destination and port; its value is a list of physical line numbers. Groups with at least two accepted entries go into the report.
Rejected rows have their own list of reasons. Two equal rejected rows do not become a repeated group. This compares accepted field values, not full original line text.
The output uses JSON, a text format for named values and lists. Groups sort by their key, while line numbers stay in input order.
4. Find repeated values
Open a command window in your folder and run:
python3 FindRepeatedEntries.py Repeated.tsv
Use your installation's Python 3 command if it is not named python3. The output is:
{
"AcceptedEntries": 3,
"RepeatedGroups": [
{
"Time": "2026-09-30T09:00:00Z",
"Source": "192.0.2.10",
"Destination": "198.51.100.20",
"Port": 443,
"Lines": [
2,
3
]
}
],
"RejectedLines": [
{
"Line": 5,
"Reason": "number outside supported range"
}
]
}
Three entries were accepted, with lines 2 and 3 in the repeated group. Line 5 is shown as rejected.
The exit code, a small result number another script can check, is 0 for no repeated groups or rejections, 1 when either needs review, and 2 for a command/load error.
5. Change the match
Save this as ChangePort.py beside the practice file:
from pathlib import Path
Log = Path("Repeated.tsv")
Data = Log.read_bytes()
Log.write_bytes(Data.replace(b"\t0443\n", b"\t80\n"))
It changes the sample's 0443 port to 80. Use it only on this practice file; it changes the saved bytes. Run:
python3 ChangePort.py
python3 FindRepeatedEntries.py Repeated.tsv
The output is now:
{
"AcceptedEntries": 3,
"RepeatedGroups": [],
"RejectedLines": [
{
"Line": 5,
"Reason": "number outside supported range"
}
]
}
The repeated group disappears because the parsed ports differ, but the exit code stays 1: line 5 is still rejected. An empty RepeatedGroups list does not mean the entire file passed. Run MakeRepeated.py again to restore the original sample. In that restored file, changing 0443 to 443 would keep the group, because both spellings become the same whole number.
Do not delete a row just because these values match. Separate actions can share them, especially with timestamps recorded only to whole seconds. You need the logging system's event identifiers and collection behaviour to decide whether a repeat is a duplicate.
The reviewer passed 21 checks, including changes to each key field, line numbering, rejections, leading-zero ports and sorted groups. Sample output was reproduced and tested inputs remained unchanged.
Save reports under a new filename: redirecting output over an input can empty it before Python opens it. Keep reports private when they reveal personal or work activity.
References
More free code guides
- Building a Time-Window Log Summary in Python
- Comparing Connection-Log Summaries in Python
- Comparing File Contents in Python
- Building a File Fingerprint in Python
- Listing a Folder's Files in Python
- Comparing Folder File Lists in Python
- Checking a Saved File Fingerprint in Python
- Finding Time Gaps in a Log with Python
- Counting Log Entries by Minute in Python