{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Did earthquakes increase on August 5?\n",
        "\n",
        "**A 30-minute data activity for grades 9-12.** Use a real, fixed snapshot of 155 earthquake records from the [USGS Comprehensive Earthquake Catalog](https://earthquake.usgs.gov/data/comcat/). The query selected worldwide records with magnitude **4.5 or greater** from **August 1 through August 7, 2026 UTC**. The CSV was downloaded on September 25, 2026. Keep `earthquakes.csv` next to this notebook and run cells from top to bottom. No account, network request, student data, or paid software is needed after download.\n",
        "\n",
        "This is a snapshot of **cataloged events meeting a selection rule**, not a count of every earthquake on Earth. USGS says catalog coverage varies by place and magnitude, and records can be updated. We will make a claim only as narrow as the data support.\n",
        "\n",
        "JR, who built Telemetry, prepared this independent activity with AI assistance. It has not been classroom-tested. It does not require Telemetry."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 1. Open the data and decide what one row means\n",
        "\n",
        "The file is a CSV table. `id` identifies a catalog event, `time` is UTC, `mag` is magnitude, and `depth` is in kilometers. Look at the first three records. Then compare the total row count with the count of unique IDs. Do they match?"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "import csv\n",
        "from pathlib import Path\n",
        "\n",
        "with Path(\"earthquakes.csv\").open(encoding=\"utf-8\", newline=\"\") as file:\n",
        "    events = list(csv.DictReader(file))\n",
        "\n",
        "print(\"Rows:\", len(events))\n",
        "print(\"Unique event IDs:\", len({event[\"id\"] for event in events}))\n",
        "for event in events[:3]:\n",
        "    print(event[\"time\"], event[\"mag\"], event[\"depth\"], event[\"place\"])\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 2. Count cataloged events by UTC day\n",
        "\n",
        "The date is the first ten characters of `time`. Make a simple text bar chart. Which day has the largest count? Describe the result as **records in this selected catalog snapshot**, not as all earthquakes worldwide."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "from collections import Counter\n",
        "\n",
        "daily = Counter(event[\"time\"][:10] for event in events)\n",
        "for day in sorted(daily):\n",
        "    print(day, f\"{daily[day]:2d}\", \"#\" * daily[day])\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 3. Change the selection rule\n",
        "\n",
        "Keep the same CSV but count only rows with `mag >= 5.0`. How many records remain? Is August 5 still the largest day? Why would it be wrong to compare the count from this stricter filter with the count from the original filter as though they measured the same thing?"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "larger = [event for event in events if float(event[\"mag\"]) >= 5.0]\n",
        "larger_daily = Counter(event[\"time\"][:10] for event in larger)\n",
        "print(\"Magnitude 5.0+ records:\", len(larger))\n",
        "for day in sorted(daily):\n",
        "    print(day, larger_daily[day])\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 4. Compare two descriptions of depth\n",
        "\n",
        "Calculate mean and median depth. Why do they differ? Inspect the largest few depths, then write one sentence that does not hide the difference. Depth is a catalog estimate; this short exercise does not assess hazard."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "from statistics import mean, median\n",
        "\n",
        "depths = [float(event[\"depth\"]) for event in events]\n",
        "print(\"Mean depth (km):\", round(mean(depths), 1))\n",
        "print(\"Median depth (km):\", round(median(depths), 1))\n",
        "print(\"Largest five depths (km):\", sorted(depths, reverse=True)[:5])\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 5. Write a claim with its boundary\n",
        "\n",
        "Finish this sentence: “In the downloaded USGS catalog snapshot for August 1-7, 2026, among records with magnitude 4.5 or greater, ______. This alone does **not** show ______.”\n",
        "\n",
        "Discuss: What would you need before claiming earthquakes are becoming more common over time? Consider observation duration, magnitude threshold, geographic coverage, catalog updates, and how networks detect smaller events.\n",
        "\n",
        "**Source and reuse:** [USGS ComCat documentation](https://earthquake.usgs.gov/data/comcat/) and [FDSN event API](https://earthquake.usgs.gov/fdsnws/event/1/1). Query: `format=csv&starttime=2026-08-01&endtime=2026-08-08&minmagnitude=4.5&orderby=time-asc`. The included CSV is the September 25 snapshot. No USGS endorsement is implied. For reuse of this notebook, credit JR / Telemetry and the USGS data source. This AI-assisted activity is untested; teachers should review it before use."
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "name": "python",
      "version": "3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 5
}
