Skip to content

Regional extractors🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

RegionalExtractor 🔗

Bases: BaseExtractor

Extracts regional-scale time series analytics for administrative boundaries.

Retrieves aggregated vegetation and crop monitoring indices at regional level (countries, states, municipalities). Supports multiple indicator types and historical gap-filling.

Documentation: https://docs.earthdaily.com/agro/library/Regional_Monitoring/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_VTS.ipynb

Args (setup_regional_parameters): index (str): Regional index type. Default: 'vegetation-vigor-index' start_date (str): Start date in YYYY-MM-DD format. Default: '2025-01-01' end_date (str): End date in YYYY-MM-DD format. Default: None (last day of current year) fillyeargap (bool): Fill year gaps in data. Default: False idblock (str): Block identifier. Default: None idpixeltype (str): Pixel type identifier. Default: None indicatorTypeIds (list): Indicator type IDs — accepted values are 1-5 only (1 VVI, 2/3 weather observed, 4/5 weather forecast). Default: None (auto-set to [1] when index='vegetation-vigor-index') use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): amu_id (required — regional AMU entity ID)

Output columns

entity_id, date, year, regional index values

get_new_token 🔗

get_new_token()

Implements token refresh logic for RegionalExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_regional_parameters 🔗

setup_regional_parameters(index='vegetation-vigor-index', start_date='2025-01-01', end_date=None, fillyeargap=False, idblock=None, idpixeltype=None, indicatorTypeIds=None, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure params for regional analytic extraction for an entity id.

Parameters:

Name Type Description Default
index

Type of vegetation index (default: "vegetation-vigor-index")

'vegetation-vigor-index'
start_date

Start date in YYYY-MM-DD format

'2025-01-01'
end_date

End date in YYYY-MM-DD format (default: last day of current year)

None
fillyeargap

Whether to fill year gaps in data

False
idblock

Block ID (267 or 141)

None
idpixeltype

Pixel type ID - [1] ALL_VEGETATIONS - [2] SUMMER - [3] WINTER - [7] GRASSLANDS - [8] ALL CROPS - [10] CORN - [11] SOYBEAN - [301] BR SOYBEAN - [400] EUR CORN - [401] EUR WINTER WHEAT - [406] EUR BARLEY - [407] EUR RAPESEED - [402] EUR SURGARBEET

None
indicatorTypeIds

List of indicator type IDs. Only 1-5 are accepted; any other value raises ValueError: - [1] for VVI - [2] for WEATHER OBSERVED ECMWF all blocks - [3] for WEATHER OBSERVED AROME on over France - [4] for WEATHER FORECAST ECMWF all blocks - [5] for WEATHER FORECAST GFS all blocks WEATHER REANALYSIS (10) and WEATHER RAINFALL ESTIMATES HI-RES / CHIRPS (11) are documented by the API but are not in this extractor's accepted set — widen valid_indicatorTypeIds below before passing them.

None
partial_frequency

Frequency threshold for partial data

50

safe_convert_amu_id staticmethod 🔗

safe_convert_amu_id(amu_id: Union[str, int, float]) -> int

Safely convert amu_id to integer, handling various input formats.

Handles: - Already an int: returns as-is - String integer: "141" -> 141 - String tuple representation: '(141;"County";...)' -> 141 - Float: 141.0 -> 141

Parameters:

Name Type Description Default
amu_id Union[str, int, float]

The AMU ID in any format

required

Returns:

Name Type Description
int int

Cleaned integer AMU ID

Raises:

Type Description
ValueError

If amu_id cannot be converted to int

Examples:

>>> safe_convert_amu_id(141)
141
>>> safe_convert_amu_id("141")
141
>>> safe_convert_amu_id('(141;"County";"Missouri")')
141
>>> safe_convert_amu_id(141.0)
141

get_regional_ts_by_id 🔗

get_regional_ts_by_id(entity_data: dict)

Query the regional analytic API to get cumulative vegetation vigor data for a given entity id.

Parameters:

Name Type Description Default
entity_data dict

Must contain the mapped 'amu_id' column (configurable via column_mapping).

required

Returns:

Name Type Description
dict

Raw JSON response from API.

get_regional_ts_by_id_safe 🔗

get_regional_ts_by_id_safe(entity_data: dict) -> dict

Safe wrapper around get_regional_ts_by_id(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain the mapped 'amu_id' column (configurable via column_mapping).

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_regional_json 🔗

format_regional_json(response_json, entity_id='unknown', base_year=1900)

Process API time series response to extract daily and average data.

Parameters:

Name Type Description Default
response_json str | dict

API response containing time series data

required
entity_id str

Entity ID for logging purposes

'unknown'
base_year int

Base year for converting dayOfYear codes (default 1900)

1900

Returns:

Name Type Description
tuple

(observed_df, daily_avg_df) - Two DataFrames with formatted data

process_single_entity_regional 🔗

process_single_entity_regional(row, params=None)

Process regional extraction for a single region with retry logic.

Parameters:

Name Type Description Default
row dict | Series

Entity data containing the mapped 'amu_id' column

required
params dict

regional parameters

None

Returns:

Name Type Description
dict

{ "observed_data": DataFrame or None, # Time series observations "daily_avg_data": DataFrame or None, # Daily averages (climatology) "error": dict or None

}

process_entity_regional_bulk_parallel 🔗

process_entity_regional_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='regional', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity regional data using threads + progress bar, with optional fail-safe retry and filter capabilities. Returns TWO normalized dataframes: observed time series and daily averages (climatology).

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain the mapped 'amu_id' column.

required
params dict

Regional parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "regional".

'regional'
use_cache bool

If True, use caching for bulk extraction. Default: None (uses instance setting).

None

Returns:

Name Type Description
dict

Contains TWO results DataFrames (observed_data and daily_avg_data), errors list, and summary statistics.