Counting Log Entries by Minute in Python
Group accepted log entries into calendar minutes and print the counts in time order. Rejected rows stay visible so an absent minute is not mistaken for proof of no activity.
You'll need Python 3, a plain-text editor and a command window. No extra packages are required.
1. Create the log file
Use a tab-separated text file with four columns: time, source address, destination address and destination port. The header must be Time, Source, Destination, Port, with actual tabs between the names. A .tsv file is a plain-text table that uses tabs to separate columns.
Write times as YYYY-MM-DDThh:mm:ssZ. The Z means Coordinated Universal Time (UTC), the shared reference time rather than a local clock. This reader accepts whole seconds in that exact form, not fractional seconds, local offsets such as +01:00 or leap-second values. Addresses use Internet Protocol version 4 (IPv4), four dot-separated numbers such as 192.0.2.10, without extra leading zeros. Ports are whole numbers from 1 to 65535.
The sample addresses are reserved for documentation in RFC 5737, a published document describing addresses to use in examples. This is the reader's own input format, not a direct export from a network-monitoring product.
Save this as MakeMinutes.py in a practice folder:
from pathlib import Path
Rows = [
"Time\tSource\tDestination\tPort\n",
"2026-09-30T09:02:00Z\t192.0.2.10\t198.51.100.20\t443\n",
"2026-09-30T09:00:00Z\t192.0.2.10\t198.51.100.20\t443\n",
"2026-09-30T09:00:59Z\t192.0.2.11\t198.51.100.20\t443\n",
"2026-09-30T09:01:00Z\t192.0.2.10\t198.51.100.20\t70000\n",
]
Path("Minutes.tsv").write_bytes("".join(Rows).encode("ascii"))
Run:
python3 MakeMinutes.py
Use your installation's Python 3 command if it is not named python3. This creates or replaces Minutes.tsv, so use a practice folder without a file you need to keep. The \t sequences write actual tabs and \n writes line feeds; you do not need to insert tabs in your editor.
The accepted rows are out of time order. The 09:01:00 row has an unsupported port and will not be counted.
2. Add the reader
The complete reader is included below so you can run this guide on its own.
Save this as LogReader.py. This reusable source file checks each row and keeps its physical line number, either with accepted values or a reason it could not be read.
from dataclasses import dataclass as DataClass
from datetime import datetime as DateTime
from pathlib import Path
@DataClass
class LogEntry:
Number: int
Original: str
EventTime: str = ""
Source: str = ""
Destination: str = ""
Port: int = 0
Problem: str = ""
Accepted: bool = False
def Digits(Value, Maximum):
if not Value:
raise ValueError("empty number")
if any(Character not in "0123456789" for Character in Value):
raise ValueError("number needs ordinary digits")
Number = int(Value)
if Number > Maximum:
raise ValueError("number outside supported range")
return Number
def CheckAddress(Value):
Parts = Value.split(".")
if len(Parts) != 4:
raise ValueError("expected four IPv4 address parts")
for Part in Parts:
if len(Part) > 3 or (len(Part) > 1 and Part.startswith("0")):
raise ValueError("use canonical dotted IPv4 addresses")
Digits(Part, 255)
def CheckTime(Value):
if len(Value) != 20:
raise ValueError("expected time YYYY-MM-DDThh:mm:ssZ")
if any(Value[Index] != Mark for Index, Mark in
((4, "-"), (7, "-"), (10, "T"), (13, ":"), (16, ":"), (19, "Z"))):
raise ValueError("expected time YYYY-MM-DDThh:mm:ssZ")
NumberText = Value[0:4] + Value[5:7] + Value[8:10] + Value[11:13] + Value[14:16] + Value[17:19]
if any(Character not in "0123456789" for Character in NumberText):
raise ValueError("time needs ordinary digits")
Year = Digits(Value[0:4], 9999)
Month = Digits(Value[5:7], 12)
Day = Digits(Value[8:10], 31)
Hour = Digits(Value[11:13], 23)
Minute = Digits(Value[14:16], 59)
Second = Digits(Value[17:19], 59)
try:
DateTime(Year, Month, Day)
except ValueError:
raise ValueError("date does not exist") from None
try:
DateTime(Year, Month, Day, Hour, Minute, Second)
except ValueError:
raise ValueError("time does not exist") from None
def ParseEntry(Entry):
try:
if len(Entry.Original) > 1024:
raise ValueError("line exceeds 1024 bytes")
if any(Character != "\t" and not 32 <= ord(Character) <= 126
for Character in Entry.Original):
raise ValueError("unsupported byte in practice format")
Fields = Entry.Original.split("\t")
if len(Fields) != 4:
raise ValueError("expected exactly four tab-separated fields")
CheckTime(Fields[0])
CheckAddress(Fields[1])
CheckAddress(Fields[2])
Entry.Port = Digits(Fields[3], 65535)
if Entry.Port == 0:
raise ValueError("port must be 1 to 65535 in this format")
Entry.EventTime, Entry.Source, Entry.Destination = Fields[:3]
Entry.Accepted = True
except ValueError as Error:
Entry.Problem = str(Error)
return Entry
def ReadLog(FileName):
with Path(FileName).open("rb") as Input:
Data = Input.read(1048577)
if len(Data) > 1048576:
raise ValueError("input exceeds 1 MiB teaching limit")
if not Data:
raise ValueError("empty input")
Lines = Data.split(b"\n")
if Lines[-1] == b"":
Lines.pop()
Lines = [Line[:-1] if Line.endswith(b"\r") else Line for Line in Lines]
if Lines[0] != b"Time\tSource\tDestination\tPort":
raise ValueError("unsupported header")
if len(Lines) - 1 > 1000:
raise ValueError("more than 1000 data lines")
return [ParseEntry(LogEntry(Number, Line.decode("latin-1")))
for Number, Line in enumerate(Lines[1:], start=2)]
The reader handles up to 1,000 data lines in a file of at most 1 MiB, roughly one million bytes, and rejects data lines longer than 1,024 bytes. Its fields use ordinary printable English-character bytes and tabs. Use a saved file that will not change while it is read.
3. Add the minute counter
Save this as CountByMinute.py:
import json as Json
import sys as Sys
from LogReader import ReadLog
def CountByMinute(FilePath):
Counts = {}
Rejected = []
Accepted = 0
for Entry in ReadLog(FilePath):
if not Entry.Accepted:
Rejected.append({"Line": Entry.Number, "Reason": Entry.Problem})
continue
Accepted += 1
Minute = Entry.EventTime[:16] + ":00Z"
Counts[Minute] = Counts.get(Minute, 0) + 1
Minutes = [{"MinuteStart": Minute, "Entries": Counts[Minute]} for Minute in sorted(Counts)]
return {"AcceptedEntries": Accepted, "ObservedMinutes": Minutes, "RejectedLines": Rejected}
def Main():
if len(Sys.argv) != 2:
print("Usage: python3 CountByMinute.py input.tsv", file=Sys.stderr)
return 2
try:
Report = CountByMinute(Sys.argv[1])
except (OSError, ValueError) as Problem:
print(f"Count stopped: {Problem}", file=Sys.stderr)
return 2
print(Json.dumps(Report, ensure_ascii=True, indent=2))
return 1 if Report["RejectedLines"] else 0
if __name__ == "__main__":
Sys.exit(Main())
Entry.EventTime[:16] keeps the checked timestamp's date, hour and minute. Adding :00Z labels that minute's start. The fixed format matters; this is not a way to round any date string.
Counts is a dictionary, a collection for looking up values by keys. Each minute-start string points to a count. Seconds 00 through 59 join the same minute; 09:01:00 starts the next one.
Sources, destinations and ports are combined, and repeated accepted rows count again. Minute labels are sorted before printing JSON, a text format for named values and lists. Rejected rows have their own list of reasons.
4. Read the minute counts
Open a command window in the folder and run:
python3 CountByMinute.py Minutes.tsv
Use your installation's Python 3 command if it is not named python3. The output is:
{
"AcceptedEntries": 3,
"ObservedMinutes": [
{
"MinuteStart": "2026-09-30T09:00:00Z",
"Entries": 2
},
{
"MinuteStart": "2026-09-30T09:02:00Z",
"Entries": 1
}
],
"RejectedLines": [
{
"Line": 5,
"Reason": "number outside supported range"
}
]
}
Two accepted entries fall in 09:00, and one in 09:02. There is no 09:01 record under ObservedMinutes because no accepted row was counted there. The rejected row still appears as line 5.
The exit code, a small result number another script can check, is 0 for a completed load without rejected rows, 1 with rejections, and 2 for a command/load error. A high count does not change it; the program has no unusual-activity threshold.
5. Fill the missing minute
Save this as AcceptMinute.py beside the practice file:
from pathlib import Path
Log = Path("Minutes.tsv")
Data = Log.read_bytes()
Log.write_bytes(Data.replace(b"\t70000\n", b"\t443\n"))
It replaces the sample's invalid port with 443. Use it only on this practice file; it changes the saved bytes. Run:
python3 AcceptMinute.py
python3 CountByMinute.py Minutes.tsv
The output is now:
{
"AcceptedEntries": 4,
"ObservedMinutes": [
{
"MinuteStart": "2026-09-30T09:00:00Z",
"Entries": 2
},
{
"MinuteStart": "2026-09-30T09:01:00Z",
"Entries": 1
},
{
"MinuteStart": "2026-09-30T09:02:00Z",
"Entries": 1
}
],
"RejectedLines": []
}
The fourth row is accepted, so 09:01:00 appears with a count of one. Its earlier absence came from rejection, not proof of zero activity. RejectedLines is empty and the exit code is now 0. If you run MakeMinutes.py again, it restores the original sample with the invalid port and the earlier exit code 1.
Only minutes containing accepted rows appear. These are calendar-minute counts, not rates for complete traffic: collection can cover part of the first or last minute, and the file does not declare its collection boundaries. Repeated rows are not automatically distinct actions, and sorting does not fix inaccurate source clocks.
The counter passed 20 checks, including minute boundaries, midnight, ordering, repeats, omitted minutes and rejected input. Sample output was reproduced and tested inputs remained unchanged.
Save reports under a new filename: redirecting output over an input can empty it before Python opens it. Keep reports private when they reveal personal or work activity.
References
More free code guides
- Building a Time-Window Log Summary in Python
- Comparing Connection-Log Summaries in Python
- Comparing File Contents in Python
- Building a File Fingerprint in Python
- Listing a Folder's Files in Python
- Comparing Folder File Lists in Python
- Checking a Saved File Fingerprint in Python
- Finding Repeated Log Entries in Python
- Finding Time Gaps in a Log with Python