The goal of these slides are to:
Introduce the modern USGS water data concepts
Introduce basic dataretrieval workflows (R and Python)
Introduce additional topics that often come up
dataRetrieval is available on the Comprehensive R Archive Network (CRAN) repository. To install dataRetrieval on your computer, open RStudio and run this line of code in the Console:
Then each time you open R, you’ll need to load the library:
R package: https://doi-usgs.github.io/dataRetrieval/
Python package: https://doi-usgs.github.io/dataretrieval-python/
WDFN Blog: https://waterdata.usgs.gov/blog/
Within R, you can call help files for any dataRetrieval function:
Within Python, you can call help for any dataretrieval function:
Help on function get_daily in module dataretrieval.waterdata.time_series:
get_daily(monitoring_location_id: 'str | Iterable[str] | None' = None, parameter_code: 'str | Iterable[str] | None' = None, statistic_id: 'str | Iterable[str] | None' = None, properties: 'str | Iterable[str] | None' = None, time_series_id: 'str | Iterable[str] | None' = None, daily_id: 'str | Iterable[str] | None' = None, approval_status: 'str | Iterable[str] | None' = None, unit_of_measure: 'str | Iterable[str] | None' = None, qualifier: 'str | Iterable[str] | None' = None, value: 'str | Iterable[str] | None' = None, last_modified: 'str | Iterable[str] | None' = None, skip_geometry: 'bool | None' = None, time: 'str | Iterable[str] | None' = None, bbox: 'list[float] | None' = None, limit: 'int | None' = None, filter: 'str | None' = None, filter_lang: 'FILTER_LANG | None' = None, convert_type: 'bool' = True, max_rows: 'int | None' = None, **queryables: 'Any') -> 'tuple[pd.DataFrame, BaseMetadata]'
Get daily values: one value per monitoring location, parameter, and day.
Throughout much of the history of the USGS, the primary water data available
was daily data collected manually at the monitoring location once each day.
With improved availability of computer storage and automated transmission of
data, the daily data published today are generally a statistical summary or
metric of the continuous data collected each day, such as the daily mean,
minimum, or maximum value. Daily data are automatically calculated from the
continuous data of the same parameter code and are described by parameter
code and a statistic code. These data have also been referred to as “daily
values” or “DV”.
Parameters
----------
monitoring_location_id : string or iterable of strings, optional
A unique identifier representing a single monitoring location,
corresponding to the id field in the monitoring-locations endpoint. IDs
combine the agency code of the agency responsible for the monitoring
location (e.g. USGS) with the location's ID number (e.g. 02238500),
separated by a hyphen (e.g. USGS-02238500).
parameter_code : string or iterable of strings, optional
A 5-digit code identifying the constituent measured and the units of
measure. A complete list of parameter codes and associated groupings is
available at https://help.waterdata.usgs.gov/codes-and-parameters/parameters.
statistic_id : string or iterable of strings, optional
A code corresponding to the statistic an observation represents.
Example codes include 00001 (max), 00002 (min), and 00003 (mean).
A complete list of codes and their descriptions can be found at
https://help.waterdata.usgs.gov/code/stat_cd_nm_query?stat_nm_cd=%25&fmt=html.
properties : string or iterable of strings, optional
The columns to return from the query.
Available options are: geometry, id, time_series_id,
monitoring_location_id, parameter_code, statistic_id, time, value,
unit_of_measure, approval_status, qualifier, last_modified
time_series_id : string or iterable of strings, optional
A unique identifier representing a single time series, corresponding to
the id field in the time-series-metadata endpoint.
daily_id : string or iterable of strings, optional
A universally unique identifier (UUID) representing a single version of
a record. The UUID is not stable over time: every time the record is
refreshed in our database, a new ID is generated. A refresh may happen
as part of normal operations and does not imply any change to the data
itself. To uniquely identify a single observation over time, compare the
time and time_series_id fields; each time series has only a single
observation at a given time.
approval_status : string or iterable of strings, optional
The approval status of each record: either "Approved", meaning
processing review has been completed and the data are approved for
publication, or "Provisional", meaning the data are subject to revision.
Some of the data you obtain from this U.S. Geological Survey database
may not have received Director's approval. Any such data values are
qualified as provisional and are subject to revision. Provisional data
are released on the condition that neither the USGS nor the United
States Government may be held liable for any damages resulting from
their use. For more information about provisional data, see
https://waterdata.usgs.gov/provisional-data-statement/.
unit_of_measure : string or iterable of strings, optional
A human-readable description of the units of measurement associated
with an observation.
qualifier : string or iterable of strings, optional
Any qualifiers associated with an observation, for instance whether a
sensor may have been impacted by ice or whether values were estimated.
value : string or iterable of strings, optional
The value of the observation. Values are transmitted as strings in
the JSON response format to preserve precision.
last_modified : string, optional
The last time a record was refreshed in our database. A refresh may
happen due to regular operational processes and does not necessarily
indicate that anything about the measurement has changed. You can query
this field using date-times or intervals, adhering to RFC 3339, or using
ISO 8601 duration objects. Intervals may be bounded or half-bounded
(double-dots at start or end). Only features whose last_modified
intersects the requested value are selected.
Examples:
* A date-time: "2018-02-12T23:20:50Z"
* A bounded interval: "2018-02-12T00:00:00Z/2018-03-18T12:31:12Z"
* Half-bounded intervals: "2018-02-12T00:00:00Z/.." or
"../2018-03-18T12:31:12Z"
* Duration objects: "P1M" for data from the past month or
"PT36H" for the last 36 hours
skip_geometry : boolean, optional
If True, the response omits the geometry of each feature and the
returned object is a data frame with no spatial information. The USGS
Water Data APIs use camelCase "skipGeometry" in CQL2 queries.
time : string, optional
The date an observation represents. You can query this field using
date-times or intervals, adhering to RFC 3339, or using ISO 8601
duration objects. Intervals may be bounded or half-bounded (double-dots
at start or end). Only features whose time intersects the requested
value are selected. If a feature has multiple temporal properties, the
server decides whether to use a single property or all relevant ones to
determine the extent.
Examples:
* A date-time: "2018-02-12T23:20:50Z"
* A bounded interval: "2018-02-12T00:00:00Z/2018-03-18T12:31:12Z"
* Half-bounded intervals: "2018-02-12T00:00:00Z/.." or
"../2018-03-18T12:31:12Z"
* Duration objects: "P1M" for data from the past month or
"PT36H" for the last 36 hours
bbox : list of numbers, optional
Only features whose geometry intersects the bounding box are selected.
The bounding box is provided as four or six numbers, depending on
whether the coordinate reference system includes a vertical axis (height
or depth). Coordinates are assumed to be in crs 4326. The expected
format is ``[xmin, ymin, xmax, ymax]``, i.e. ``[Western-most longitude,
Southern-most latitude, Eastern-most longitude, Northern-most
latitude]``.
limit : int, optional
The number of features returned in each page. The maximum allowable
limit is 50000; the default (None) requests that maximum. Set a lower
number if your internet connection is spotty. This is a per-page size,
not a cap on the total result: a query matching more rows than ``limit``
still returns every matching row across multiple pages. Use ``max_rows``
to cap the total instead.
filter, filter_lang : optional
Server-side CQL filter passed through as the OGC ``filter`` /
``filter-lang`` query parameters. See
:mod:`dataretrieval.ogc.filters` for syntax, auto-chunking,
and the lexicographic-comparison pitfall.
convert_type : boolean, optional
If True, converts columns to appropriate types.
max_rows : int, optional
Cap the total number of rows returned, stopping pagination early
instead of downloading the whole result. Unlike ``limit`` (the
per-page size), this bounds the total result across every page.
The default (None) follows pagination to completion.
**queryables : string or iterable of strings, optional
Any other queryable property of this collection, passed through as a
server-side filter. Call :func:`get_queryables` to see the queryables a
collection supports.
Returns
-------
df : ``pandas.DataFrame`` or ``geopandas.GeoDataFrame``
Formatted data returned from the API query.
md: :obj:`dataretrieval.utils.BaseMetadata`
A custom metadata object
Raises
------
ChunkInterrupted
A transient failure (429 / 5xx / timeout) interrupted the request
after the built-in retries. Completed work is preserved; resume
with ``exc.call.resume()`` (see :doc:`/userguide/errors`).
Examples
--------
.. code::
>>> # Get daily flow data from a single site
>>> # over a yearlong period
>>> df, md = dataretrieval.waterdata.get_daily(
... monitoring_location_id="USGS-02238500",
... parameter_code="00060",
... time="2021-01-01T00:00:00Z/2022-01-01T00:00:00Z",
... )
>>> # Quick "show me the last week" idiom (ISO 8601 duration)
>>> df, md = dataretrieval.waterdata.get_daily(
... monitoring_location_id="USGS-02238500",
... parameter_code="00060",
... time="P7D",
... )
>>> # Get approved daily flow data from multiple sites
>>> df, md = dataretrieval.waterdata.get_daily(
... monitoring_location_id=["USGS-05114000", "USGS-09423350"],
... approval_status="Approved",
... time="2024-01-01/..",
... )
>>> # Pull only rows whose underlying record was refreshed in the
>>> # last 7 days — handy for incremental ETL polling
>>> df, md = dataretrieval.waterdata.get_daily(
... monitoring_location_id="USGS-02238500",
... parameter_code="00060",
... last_modified="P7D",
... )
>>> # Chain queries: pull all stream sites in a state, then their
>>> # daily discharge for the last week. The site list can be hundreds
>>> # of values long — the request is transparently chunked across
>>> # multiple chunks so the URL stays under the server's byte
>>> # limit. Combined output looks like a single query.
>>> sites_df, _ = dataretrieval.waterdata.get_monitoring_locations(
... state="Ohio",
... site_type="Stream",
... )
>>> df, md = dataretrieval.waterdata.get_daily(
... monitoring_location_id=sites_df["monitoring_location_id"].tolist(),
... parameter_code="00060",
... time="P7D",
... )
USGS Water Data APIs
Continuous (e.g. 15-minute sensor data)
Daily (e.g. mean from continuous)
Monitoring Location Information
Time Series Information
Latest Daily/Continuous
Field Measurements
USGS Water Data APIs
Peak Flows
Rating Curves
Discrete Water-Quality
Water Quality Portal (WQP) Data
Each of these is has a different:
dataRetrieval functionThe USGS uses various codes for basic retrievals. These codes can have leading zeros, therefore they need to be a character surrounded in quotes (“00060”).
read_metadata("parameter-codes")read_metadata("statistic-codes")Here are some examples of a few common codes:
|
|
dataRetrieval can help!parameter_codes <- read_waterdata_metadata("parameter-codes")
statistic_codes <- read_waterdata_metadata("statistic-codes")
# Others:
agency_codes <- read_waterdata_metadata("agency-codes")
aquifer_codes <- read_waterdata_metadata("aquifer-codes")
aquifer_types <- read_waterdata_metadata("aquifer-types")
coordinate_datum_codes <- read_waterdata_metadata("coordinate-datum-codes")
huc_codes <- read_waterdata_metadata("hydrologic-unit-codes")
national_aquifer_codes <- read_waterdata_metadata("national-aquifer-codes")
reliability_codes <- read_waterdata_metadata("reliability-codes")
site_types <- read_waterdata_metadata("site-types")
topographic_codes <- read_waterdata_metadata("topographic-codes")
time_zone_codes <- read_waterdata_metadata("time-zone-codes")
counties <- read_waterdata_metadata("counties")
states <- read_waterdata_metadata("states")parameter_codes = waterdata.get_reference_table("parameter-codes")
statistic_codes = waterdata.get_reference_table("statistic-codes")
# Others:
agency_codes = waterdata.get_reference_table("agency-codes")
aquifer_codes = waterdata.get_reference_table("aquifer-codes")
aquifer_types = waterdata.get_reference_table("aquifer-types")
coordinate_datum_codes = waterdata.get_reference_table("coordinate-datum-codes")
huc_codes = waterdata.get_reference_table("hydrologic-unit-codes")
national_aquifer_codes = waterdata.get_reference_table("national-aquifer-codes")
reliability_codes = waterdata.get_reference_table("reliability-codes")
site_types = waterdata.get_reference_tablea("site-types")
topographic_codes = waterdata.get_reference_table("topographic-codes")
time_zone_codes = waterdata.get_reference_table("time-zone-codes")
counties = waterdata.get_reference_table("counties")
states = waterdata.get_reference_table("states")Each function returns a Tuple, containing a dataframe and a Metadata class.
Open your preferred IDE (RStudio, VSCode, PyCharm, etc) or Jupyter notebook
Install packages if needed:
R: dataRetrieval, ggplot2, dplyr, leaflet
Python: dataretrieval, matplotlib, geopandas, folium, seaborn
Load dataRetrieval (R) / waterdata module in dataretrieval (Python)
Open the help file for the function read_waterdata_daily (R) or waterdata.get_daily (Python)
The Water Data APIs limit how many queries a single IP address can make per hour
You can run new dataRetrieval functions without a token
You might run into errors quickly. If you (or your IP!) have exceeded the quota, you will see:
! HTTP 429 Too Many Requests.
• You have exceeded your rate limit. Make sure you provided your API key from https://api.waterdata.usgs.gov/signup/, then either try again later or contact us at https://waterdata.usgs.gov/questions-comments/?referrerUrl=https://api.waterdata.usgs.gov for assistance.
Request a USGS Water Data API Token: https://api.waterdata.usgs.gov/signup/
Save it in a safe place (KeyPass or other password management tool)
Add it as environment variable
Restart
See next slide for a demonstration.
Let’s pretend the token sent you was “abc123”
API_USGS_PAT = "abc123"
Save that file
Restart R/RStudio.
Check that it worked by running (you should see your token printed in the Console):
Note
Your .Renviron file should never be pushed to a public repository.
Create a file in your working directory .env
Add this line to the .env file:
API_USGS_PAT = "abc123"
Restart your python session
Check that it worked by running (you should see your token printed in the Console):
Note
Your .env file should never be pushed to a public repository.
Open Miniforge, Anaconda, etc.
Activate enviornment
Use your “tab” key!


Shift + Tab:

When you look at the help file for the new functions, you’ll notice there are lots of possible inputs parameters.
You DO NOT need to (and should not!) specify all of these parameters.
However, also consider what happens if you leave too many things blank. What do you suppose will be returned here?
Since no list of sites or bounding box was defined, ALL the daily data in ALL the country with parameter code “00060” and statistic code “00003” will be returned.
Time parameters have a few options:
A single date (or date-time): “2024-10-01” or “2024-10-01T23:20:50Z”
A bounded interval: c(“2024-10-01”, “2025-07-02”)
Half-bounded intervals: c(“2024-10-01”, NA)
Duration objects: “P1M” for data from the past month or “PT36H” for the last 36 hours
Here are a bunch of valid inputs:
# Ask for exact times:
time = "2025-01-01"
time = as.Date("2025-01-01")
time = "2025-01-01T23:20:50Z"
time = as.POSIXct(
"2025-01-01T23:20:50Z",
format = "%Y-%m-%dT%H:%M:%S",
tz = "UTC"
)
# Ask for specific range
time = c("2024-01-01", "2025-01-01") # or Dates or POSIXs
# Asking beginning of record to specific end:
time = c(NA, "2024-01-01") # or Date or POSIX
# Asking specific beginning to end of record:
time = c("2024-01-01", NA) # or Date or POSIX
# Ask for period
time = "P1M" # past month
time = "P7D" # past 7 days
time = "PT12H" # past hours# Ask for exact times:
time = "2025-01-01"
# Ask for specific range
time = "2025-01-01/2026-01-01"
# Asking beginning of record to specific end:
time = "../2024-01-01" # or Date or POSIX
# Asking specific beginning to end of record:
time = "2024-01-01/.." # or Date or POSIX
# Ask for period
time = "P1M" # past month
time = "P7D" # past 7 days
time = "PT12H" # past hoursWorkflow 1: Find Available Sites
Workflow 2: Find Available Data
Challenge 1
Workflow 3: Get Latest Data
Workflow 4: Get All Data
Challenge 2
Workflow 5: Discrete Water Quality
Let’s get all the monitoring locations for Dane County, Wisconsin:
Now that we’ve seen the whole data set, maybe we realize in the future we can ask for just stream sites, and we only really need a few of those columns:
Requesting:
https://api.waterdata.usgs.gov/ogcapi/v1/collections/monitoring-locations/items?f=json&lang=en-US&properties=monitoring_location_name%2Cdrainage_area&state_name=Wisconsin&county_name=Dane%20County&site_type=Stream&limit=50000
Remaining requests this hour:933
If you have geopandas installed, the function will return a GeoDataFrame with a geometry column containing the monitoring locations’ coordinates. You can use gpd.explore() to view your geometry coordinates on an interactive map.
Let’s get all the time series in Dane County, WI with daily mean (statistic_id = “00003”) discharge (parameter code = “00060”) or temperature (parameter code = “00010”).
Selecting just a few columns:
How many USGS stream sites are within Tuscaloosa County, Alabama?
What are the unique parameter_names that come back from those stream sites?
Of those sites, what site has the longest period of record for daily mean discharge?
Bonus: Create an interactive map of all sites that measure daily mean discharge.
The amount you get done during this break will highly depend on the extent of your coding background. Use this time to explore dataretrieval functions and outputs.
When these slides were generated on 2026-09-09, the results were:
1.
[1] 99
2.
[1] "Gage height" "Discharge" "Suspnd sedmnt disch"
[4] "Precipitation, cumul" "Suspnd sedmnt conc" "Precipitation"
3.
[1] "BLACK WARRIOR RIVER AT NORTHPORT AL"
[1] "1895-01-01 06:00:00 UTC"
[1] "2026-09-07 05:00:00 UTC"
Let’s get the continuous discharge measurements in Dane County, WI (parameter code = “00060”) that have measured data within the last 14 days.
latest_sites <- read_waterdata_combined_meta(
state_name = "Wisconsin",
county_name = "Dane County",
parameter_code = c("00060"),
last_modified = "P14D",
data_type = "Continuous values"
)
latest_discharge <- read_waterdata_latest_continuous(
monitoring_location_id = latest_sites$monitoring_location_id,
parameter_code = "00060"
)latest_sites, md = waterdata.get_combined_metadata(
state_name = "Wisconsin",
county_name = "Dane County",
parameter_code = "00060",
last_modified = "P14D",
data_type = "Continuous values"
)
latest_discharge, md = waterdata.get_latest_continuous(
monitoring_location_id = latest_sites.monitoring_location_id,
parameter_code = "00060"
)pal <- colorNumeric("viridis", latest_discharge$value)
leaflet(
data = latest_discharge |>
sf::st_transform(crs = leaflet_crs)
) |>
addProviderTiles("CartoDB.Positron") |>
addCircleMarkers(
popup = paste(
latest_discharge$monitoring_location_id,
"<br>",
latest_discharge$time,
"<br>",
latest_discharge$value,
latest_discharge$unit_of_measure
),
color = ~ pal(value),
radius = 3,
opacity = 1
) |>
addLegend(
pal = pal,
position = "bottomleft",
title = "Latest Discharge",
values = ~value
)Let’s get daily discharge data for the last 3 years from 2 sites:
Navigate to the National Water Dashboard https://dashboard.waterdata.usgs.gov/app/nwd/en/
Zoom in and explore sections of the map that have generally higher than normal streamflow.
Zoom in and explore sections of the map that have generally lower than normal streamflow.
Pick a USGS site that is interesting to you (maybe you are a kayaker, maybe you fish, maybe it’s a local stream, maybe it’s in an extreme flood/drought).
Plot the daily mean discharge data for all time for the site you picked.
Let’s get orthophosphate (“00660”) data from the Shenandoah River at Front Royal, VA (“USGS-01631000”).
GET: https://api.waterdata.usgs.gov/samples-data/results/basicphyschem?mimeType=text%2Fcsv&monitoringLocationIdentifier=USGS-01631000&usgsPCode=00660
[1] 104
That’s a LOT of columns returned.
Any use of trade, firm, or product name is for descriptive purposes only and does not imply endorsement by the U.S. Government.
