Introduction to dataRetrieval

Laura DeCicco

Introduction

In this ~45 minute introduction, the goal is:

  • Introduce the modern dataRetrieval workflows.

  • The intended audience is someone:

    • New to dataRetrieval

    • Has some R experience

dataRetrieval: R-package for US water data

USGS Water Data APIs *

  • Surface water levels

  • Groundwater levels

  • Site metadata

  • Peak flows

  • Rating curves

  • Discrete water-quality data

Water Quality Portal (WQP) Data

  • Discrete water-quality data

  • USGS and non-USGS data

Installation

dataRetrieval is available on the Comprehensive R Archive Network (CRAN) repository. To install dataRetrieval on your computer, open RStudio and run this line of code in the Console:

install.packages("dataRetrieval")

Then each time you open R, you’ll need to load the library:

library(dataRetrieval)
pip install dataretrieval
conda install conda-forge::dataretrieval

Then each time you open Python, you’ll need to load the library:

from dataretrieval import waterdata
Geopandas not installed. Geometries will be flattened into pandas DataFrames.

dataRetrieval: External Documentation

dataRetrieval: External Documentation

dataRetrieval: External Documentation

Documentation within R: function help pages

Within R, you can call help files for any dataRetrieval function:

?read_waterdata_daily

Click here to open a new window in RStudio:

Scroll down to the “Examples” to see how each function can be run.

Examples

site <- "USGS-02238500"
dv_data_sf <- read_waterdata_daily(
  monitoring_location_id = site,
  parameter_code = "00060",
  time = c("2021-01-01", "2022-01-01")
)

Within Python, you can call help for any dataretrieval function:

help(waterdata.get_daily)
Help on function get_daily in module dataretrieval.waterdata.time_series:

get_daily(monitoring_location_id: 'str | Iterable[str] | None' = None, parameter_code: 'str | Iterable[str] | None' = None, statistic_id: 'str | Iterable[str] | None' = None, properties: 'str | Iterable[str] | None' = None, time_series_id: 'str | Iterable[str] | None' = None, daily_id: 'str | Iterable[str] | None' = None, approval_status: 'str | Iterable[str] | None' = None, unit_of_measure: 'str | Iterable[str] | None' = None, qualifier: 'str | Iterable[str] | None' = None, value: 'str | Iterable[str] | None' = None, last_modified: 'str | Iterable[str] | None' = None, skip_geometry: 'bool | None' = None, time: 'str | Iterable[str] | None' = None, bbox: 'list[float] | None' = None, limit: 'int | None' = None, filter: 'str | None' = None, filter_lang: 'FILTER_LANG | None' = None, convert_type: 'bool' = True, max_rows: 'int | None' = None, **queryables: 'Any') -> 'tuple[pd.DataFrame, BaseMetadata]'
    Get daily values: one value per monitoring location, parameter, and day.

    Throughout much of the history of the USGS, the primary water data available
    was daily data collected manually at the monitoring location once each day.
    With improved availability of computer storage and automated transmission of
    data, the daily data published today are generally a statistical summary or
    metric of the continuous data collected each day, such as the daily mean,
    minimum, or maximum value. Daily data are automatically calculated from the
    continuous data of the same parameter code and are described by parameter
    code and a statistic code. These data have also been referred to as “daily
    values” or “DV”.

    Parameters
    ----------
    monitoring_location_id : string or iterable of strings, optional
        A unique identifier representing a single monitoring location,
        corresponding to the id field in the monitoring-locations endpoint. IDs
        combine the agency code of the agency responsible for the monitoring
        location (e.g. USGS) with the location's ID number (e.g. 02238500),
        separated by a hyphen (e.g. USGS-02238500).
    parameter_code : string or iterable of strings, optional
        A 5-digit code identifying the constituent measured and the units of
        measure. A complete list of parameter codes and associated groupings is
        available at https://help.waterdata.usgs.gov/codes-and-parameters/parameters.
    statistic_id : string or iterable of strings, optional
        A code corresponding to the statistic an observation represents.
        Example codes include 00001 (max), 00002 (min), and 00003 (mean).
        A complete list of codes and their descriptions can be found at
        https://help.waterdata.usgs.gov/code/stat_cd_nm_query?stat_nm_cd=%25&fmt=html.
    properties : string or iterable of strings, optional
        The columns to return from the query.
        Available options are: geometry, id, time_series_id,
        monitoring_location_id, parameter_code, statistic_id, time, value,
        unit_of_measure, approval_status, qualifier, last_modified
    time_series_id : string or iterable of strings, optional
        A unique identifier representing a single time series, corresponding to
        the id field in the time-series-metadata endpoint.
    daily_id : string or iterable of strings, optional
        A universally unique identifier (UUID) representing a single version of
        a record. The UUID is not stable over time: every time the record is
        refreshed in our database, a new ID is generated. A refresh may happen
        as part of normal operations and does not imply any change to the data
        itself. To uniquely identify a single observation over time, compare the
        time and time_series_id fields; each time series has only a single
        observation at a given time.
    approval_status : string or iterable of strings, optional
        The approval status of each record: either "Approved", meaning
        processing review has been completed and the data are approved for
        publication, or "Provisional", meaning the data are subject to revision.
        Some of the data you obtain from this U.S. Geological Survey database
        may not have received Director's approval. Any such data values are
        qualified as provisional and are subject to revision. Provisional data
        are released on the condition that neither the USGS nor the United
        States Government may be held liable for any damages resulting from
        their use. For more information about provisional data, see
        https://waterdata.usgs.gov/provisional-data-statement/.
    unit_of_measure : string or iterable of strings, optional
        A human-readable description of the units of measurement associated
        with an observation.
    qualifier : string or iterable of strings, optional
        Any qualifiers associated with an observation, for instance whether a
        sensor may have been impacted by ice or whether values were estimated.
    value : string or iterable of strings, optional
        The value of the observation. Values are transmitted as strings in
        the JSON response format to preserve precision.
    last_modified : string, optional
        The last time a record was refreshed in our database. A refresh may
        happen due to regular operational processes and does not necessarily
        indicate that anything about the measurement has changed. You can query
        this field using date-times or intervals, adhering to RFC 3339, or using
        ISO 8601 duration objects. Intervals may be bounded or half-bounded
        (double-dots at start or end). Only features whose last_modified
        intersects the requested value are selected.
        Examples:

            * A date-time: "2018-02-12T23:20:50Z"
            * A bounded interval: "2018-02-12T00:00:00Z/2018-03-18T12:31:12Z"
            * Half-bounded intervals: "2018-02-12T00:00:00Z/.." or
                "../2018-03-18T12:31:12Z"
            * Duration objects: "P1M" for data from the past month or
                "PT36H" for the last 36 hours

    skip_geometry : boolean, optional
        If True, the response omits the geometry of each feature and the
        returned object is a data frame with no spatial information. The USGS
        Water Data APIs use camelCase "skipGeometry" in CQL2 queries.
    time : string, optional
        The date an observation represents. You can query this field using
        date-times or intervals, adhering to RFC 3339, or using ISO 8601
        duration objects. Intervals may be bounded or half-bounded (double-dots
        at start or end). Only features whose time intersects the requested
        value are selected. If a feature has multiple temporal properties, the
        server decides whether to use a single property or all relevant ones to
        determine the extent.
        Examples:

            * A date-time: "2018-02-12T23:20:50Z"
            * A bounded interval: "2018-02-12T00:00:00Z/2018-03-18T12:31:12Z"
            * Half-bounded intervals: "2018-02-12T00:00:00Z/.." or
                "../2018-03-18T12:31:12Z"
            * Duration objects: "P1M" for data from the past month or
                "PT36H" for the last 36 hours

    bbox : list of numbers, optional
        Only features whose geometry intersects the bounding box are selected.
        The bounding box is provided as four or six numbers, depending on
        whether the coordinate reference system includes a vertical axis (height
        or depth). Coordinates are assumed to be in crs 4326. The expected
        format is ``[xmin, ymin, xmax, ymax]``, i.e. ``[Western-most longitude,
        Southern-most latitude, Eastern-most longitude, Northern-most
        latitude]``.
    limit : int, optional
        The number of features returned in each page. The maximum allowable
        limit is 50000; the default (None) requests that maximum. Set a lower
        number if your internet connection is spotty. This is a per-page size,
        not a cap on the total result: a query matching more rows than ``limit``
        still returns every matching row across multiple pages. Use ``max_rows``
        to cap the total instead.
    filter, filter_lang : optional
        Server-side CQL filter passed through as the OGC ``filter`` /
        ``filter-lang`` query parameters. See
        :mod:`dataretrieval.ogc.filters` for syntax, auto-chunking,
        and the lexicographic-comparison pitfall.
    convert_type : boolean, optional
        If True, converts columns to appropriate types.
    max_rows : int, optional
        Cap the total number of rows returned, stopping pagination early
        instead of downloading the whole result. Unlike ``limit`` (the
        per-page size), this bounds the total result across every page.
        The default (None) follows pagination to completion.
    **queryables : string or iterable of strings, optional
        Any other queryable property of this collection, passed through as a
        server-side filter. Call :func:`get_queryables` to see the queryables a
        collection supports.

    Returns
    -------
    df : ``pandas.DataFrame`` or ``geopandas.GeoDataFrame``
        Formatted data returned from the API query.
    md: :obj:`dataretrieval.utils.BaseMetadata`
        A custom metadata object

    Raises
    ------
    ChunkInterrupted
        A transient failure (429 / 5xx / timeout) interrupted the request
        after the built-in retries. Completed work is preserved; resume
        with ``exc.call.resume()`` (see :doc:`/userguide/errors`).

    Examples
    --------
    .. code::

        >>> # Get daily flow data from a single site
        >>> # over a yearlong period
        >>> df, md = dataretrieval.waterdata.get_daily(
        ...     monitoring_location_id="USGS-02238500",
        ...     parameter_code="00060",
        ...     time="2021-01-01T00:00:00Z/2022-01-01T00:00:00Z",
        ... )

        >>> # Quick "show me the last week" idiom (ISO 8601 duration)
        >>> df, md = dataretrieval.waterdata.get_daily(
        ...     monitoring_location_id="USGS-02238500",
        ...     parameter_code="00060",
        ...     time="P7D",
        ... )

        >>> # Get approved daily flow data from multiple sites
        >>> df, md = dataretrieval.waterdata.get_daily(
        ...     monitoring_location_id=["USGS-05114000", "USGS-09423350"],
        ...     approval_status="Approved",
        ...     time="2024-01-01/..",
        ... )

        >>> # Pull only rows whose underlying record was refreshed in the
        >>> # last 7 days — handy for incremental ETL polling
        >>> df, md = dataretrieval.waterdata.get_daily(
        ...     monitoring_location_id="USGS-02238500",
        ...     parameter_code="00060",
        ...     last_modified="P7D",
        ... )

        >>> # Chain queries: pull all stream sites in a state, then their
        >>> # daily discharge for the last week. The site list can be hundreds
        >>> # of values long — the request is transparently chunked across
        >>> # multiple chunks so the URL stays under the server's byte
        >>> # limit. Combined output looks like a single query.
        >>> sites_df, _ = dataretrieval.waterdata.get_monitoring_locations(
        ...     state="Ohio",
        ...     site_type="Stream",
        ... )
        >>> df, md = dataretrieval.waterdata.get_daily(
        ...     monitoring_location_id=sites_df["monitoring_location_id"].tolist(),
        ...     parameter_code="00060",
        ...     time="P7D",
        ... )

dataRetrieval Updates

Are you a seasoned dataRetrieval user?

Here are resources for recent major changes:

What’s New?

There’s been a lot of changes to dataRetrieval over the past year. If you’d like to see an overview of those changes, visit: Changes to dataRetrieval

Biggest changes:

  • NWIS servers will be shut down, so all readNWIS functions will eventually stop working

  • read_waterdata functions are modern and should be used when possible

  • The “USGS Water Data APIs” are the new home for USGS data

USGS Water Data API Token

  • The Water Data APIs limit how many queries a single IP address can make per hour

  • You can run new dataRetrieval functions without a token

  • You might run into errors quickly. If you (or your IP!) have exceeded the quota, you will see:

! HTTP 429 Too Many Requests.
  • You have exceeded your rate limit. Make sure you provided your API key from https://api.waterdata.usgs.gov/signup/, then either try again later or contact us at https://waterdata.usgs.gov/questions-comments/?referrerUrl=https://api.waterdata.usgs.gov for assistance.

USGS Water Data API Token

  1. Request a USGS Water Data API Token: https://api.waterdata.usgs.gov/signup/

  2. Save it in a safe place (KeePass or other password management tool)

  3. Add it to your .Renviron file as API_USGS_PAT.

  4. Restart R

  5. Check that it worked by running (you should see your token printed in the Console):

Sys.getenv("API_USGS_PAT")

See next slide for a demonstration.

USGS Water Data API Token: Example

My favorite method to do add your token to .Renviron is to use the usethis package. Let’s pretend the token sent you was “abc123”:

  1. Run in R:
usethis::edit_r_environ()
  1. Add this line to the file that opens up:
API_USGS_PAT = "abc123"
  1. Save that file using the save button

  2. Restart R/RStudio.

  3. Run after restarting R:

Sys.getenv("API_USGS_PAT")

USGS Water Data API Token: Example

After save and restart, check that it worked by running:

Sys.getenv("API_USGS_PAT")

USGS Basic Retrievals

The USGS uses various codes for basic retrievals. These codes can have leading zeros, therefore they need to be a character surrounded in quotes (“00060”).

  • Site ID (often 8 or 15-digits)
  • Parameter Code (5 digits)
    • Full list: read_waterdata_parameter_codes()
  • Statistic Code (for daily values)
    • Full list: read_metadata("statistic-codes")

USGS Basic Retrievals Parameter and Statistic Codes

Here are some examples of a few common codes:

Parameter Code Short Name
00060 Discharge
00065 Gage Height
00010 Temperature
00400 pH
Statistic Code Short Name
00001 Maximum
00002 Minimum
00003 Mean
00008 Median

Let’s Go!

We’re going walk through 3 retrievals:

  • Workflow 1: Daily Data

    • Uses the new USGS Water Data API

    • Modern data access point going forward

  • Workflow 2: Discrete Data

    • Uses new USGS Samples Data

    • Modern data access point going forward

  • Workflow 3: Continuous Data

    • Uses the new USGS Water Data API

    • Modern data access point going forward

Workflow 1: Daily data for known site

Let’s pull daily mean discharge data for site “USGS-0940550”, getting all the data from October 10, 2025 onward.

library(dataRetrieval)
site <- "USGS-09405500"
pcode <- "00060" # Discharge
stat_cd <- "00003" # Mean
range <- c("2025-10-01", NA)

df <- read_waterdata_daily(
  monitoring_location_id = site,
  parameter_code = pcode,
  statistic_id = stat_cd,
  time = range
)
Requesting:
https://api.waterdata.usgs.gov/ogcapi/v1/collections/daily/items?f=json&lang=en-US&monitoring_location_id=USGS-09405500&parameter_code=00060&statistic_id=00003&time=2025-10-01%2F..&limit=50000
Remaining requests this hour:914 
nrow(df)
[1] 343
from dataretrieval import waterdata

site = "USGS-09405500"
pcode = "00060"  # Discharge
stat_cd = "00003"  # Mean

df, md = waterdata.get_daily(
    monitoring_location_id=site,
    parameter_code=pcode,
    statistic_id=stat_cd,
    time="2025-10-01/..",
)

df.shape[0]
343

Workflow 1: Look at Daily Data

In RStudio, click on the data frame in the upper right Environment tab to open a Viewer.

Workflow 1: Plot Daily Data

Let’s use ggplot2 to visualize the data.

library(ggplot2)

theme_set(theme_bw(base_size = 24))
update_geom_defaults("point", list(size = 3, color = "steelblue"))
options(ggplot2.discrete.colour = "viridis")
options(ggplot2.discrete.fill = "viridis")

ggplot(data = df) +
  geom_point(aes(x = time, y = value, color = approval_status))

Let’s use matplotlib to visualize the data.

import matplotlib.pyplot as plt
import pandas as pd

plt.rcParams["font.size"] = 20

levels, categories = pd.factorize(df["approval_status"])

fig, ax = plt.subplots()
scatter = ax.scatter(x=df.time, y=df.value, c=levels)
fig.legend(scatter.legend_elements()[0], categories, title="Status")

Water Data API Notes: Argument input

Use your “tab” key!

Water Data API Notes: Arguments

  • When you look at the help file for the new functions, you’ll notice there are lots of possible inputs (arguments).

  • You DO NOT need to (and should not!) specify all of these parameters.

  • However, also consider what happens if you leave too many things blank. What do you suppose will be returned here?

discharge <- read_waterdata_daily(
  parameter_code = "00060",
  statistic_id = "00003"
)

Since no list of sites or bounding box was defined, ALL the daily data in ALL the country with parameter code “00060” and statistic code “00003” will be returned.

Water Data API Notes: time input

The “time” argument has a few options:

  • A single date (or date-time): “2024-10-01” or “2024-10-01T23:20:50Z”

  • A bounded interval: c(“2024-10-01”, “2025-07-02”)

  • Half-bounded intervals: c(“2024-10-01”, NA)

  • Duration objects: “P1M” for data from the past month or “PT36H” for the last 36 hours

Here are a bunch of valid inputs:

# Ask for exact times:
time = "2025-01-01"
time = as.Date("2025-01-01")
time = "2025-01-01T23:20:50Z"
time = as.POSIXct(
  "2025-01-01T23:20:50Z",
  format = "%Y-%m-%dT%H:%M:%S",
  tz = "UTC"
)
# Ask for specific range
time = c("2024-01-01", "2025-01-01") # or Dates or POSIXs
# Asking beginning of record to specific end:
time = c(NA, "2024-01-01") # or Date or POSIX
# Asking specific beginning to end of record:
time = c("2024-01-01", NA) # or Date or POSIX
# Ask for period
time = "P1M" # past month
time = "P7D" # past 7 days
time = "PT12H" # past hours

Workflow 2: Discrete data for known site

Use your “tab” key!

Workflow 2: Discrete data for known site

Let’s get orthophosphate (“00660”) data from the Shenandoah River at Front Royal, VA (“USGS-01631000”).

site <- "USGS-01631000"
pcode <- "00660"

qw_data <- read_waterdata_samples(
  monitoringLocationIdentifier = site,
  usgsPCode = pcode,
  dataType = "results",
  dataProfile = "basicphyschem"
)
GET: https://api.waterdata.usgs.gov/samples-data/results/basicphyschem?mimeType=text%2Fcsv&monitoringLocationIdentifier=USGS-01631000&usgsPCode=00660
ncol(qw_data)
[1] 104

R generates a few POSIXct columns to combine date, time, timezone information.

site = "USGS-01631000"
pcode = "00660"

qw_data, md_qw = waterdata.get_samples(
    monitoringLocationIdentifier = site,
    usgsPCode = pcode,
    service = "results",
    profile = "basicphyschem",
)

qw_data.shape[1]
101

That’s a LOT of columns returned. We won’t look at them here, but you can use View in RStudio to explore on your own.

USGS Samples Data Notes: Data Types and Profiles

  • There are 2 arguments that dictate what kind of data is returned
    • “dataType” defines what kind of data comes back
    • “dataProfile” defines what columns from that type come back
  • There are 2 parameters that dictate what kind of data is returned
    • “service” defines what kind of data comes back
    • “profile” defines what columns from that type come back

Data Types and Profiles

Workflow 3: Continuous data for known site

  • Continuous data is the high-frequency sensor data.

  • We’ll look at Suisun Bay a Van Sickle Island NR Pittsburg CA (“USGS-11455508”), with parameter code “99133” which is Nitrate plus Nitrite.

Workflow 3: Continuous data for known site

site_id <- "USGS-11455508"
p_code_rt <- "99133"
start_date <- "2024-01-01"
end_date <- "2024-06-01"

continuous_data <- read_waterdata_continuous(
  monitoring_location_id = site_id,
  parameter_code = p_code_rt,
  time = c(start_date, end_date)
)
names(continuous_data)
 [1] "monitoring_location_id"
 [2] "parameter_code"        
 [3] "statistic_id"          
 [4] "time"                  
 [5] "value"                 
 [6] "unit_of_measure"       
 [7] "approval_status"       
 [8] "qualifier"             
 [9] "last_modified"         
[10] "time_series_id"        
site_id = "USGS-11455508"
p_code_rt = "99133"
date_range = "2024-01-01/2024-06-01"

continuous_data, md_cont = waterdata.get_continuous(
    monitoring_location_id=site_id, parameter_code=p_code_rt, time=date_range
)
Index(['time_series_id',
       'monitoring_location_id',
       'parameter_code',
       'statistic_id',
       'time', 'value',
       'unit_of_measure',
       'approval_status',
       'qualifier',
       'last_modified',
       'geometry',
       'continuous_id'],
      dtype='str')

Workflow 3: Inspect

ggplot(data = continuous_data) +
  geom_point(aes(x = time, y = value))

plt.figure()
plt.scatter(x=continuous_data.time, y=continuous_data.value)

Data Discovery

  1. read_waterdata_ts_meta discovers available daily and continuous time series

  2. read_waterdata_field_meta discovers available field measurement

  3. read_waterdata_combined_meta combines the time series and field measurement discovery. This function also provides the most flexibility with geographic queries.

  4. summarize_waterdata_samples discovers discrete data at specific monitoring locations

The next slides will demo how to use those.

Data Discovery: Time Series

ts_available <- read_waterdata_combined_meta(
  monitoring_location_id = "USGS-04183500"
)

Data Discovery: Discrete

discrete_available <- summarize_waterdata_samples(
  monitoringLocationIdentifier = "USGS-04183500"
)

characteristicUserSupplied

  • characteristicUserSupplied can be an input to read_waterdata_sample
discrete1 <- read_waterdata_samples(
  characteristicUserSupplied = "Phosphorus as phosphorus, water, unfiltered",
  monitoringLocationIdentifier = "USGS-04183500"
)
nrow(discrete1)
[1] 1300

More Information