Regional extractors🔗
Generated from the source docstrings by mkdocstrings at build time.
See the API reference overview for the other groups.
RegionalExtractor 🔗
Bases: BaseExtractor
Extracts regional-scale time series analytics for administrative boundaries.
Retrieves aggregated vegetation and crop monitoring indices at regional level (countries, states, municipalities). Supports multiple indicator types and historical gap-filling.
Documentation: https://docs.earthdaily.com/agro/library/Regional_Monitoring/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_VTS.ipynb
Args (setup_regional_parameters): index (str): Regional index type. Default: 'vegetation-vigor-index' start_date (str): Start date in YYYY-MM-DD format. Default: '2025-01-01' end_date (str): End date in YYYY-MM-DD format. Default: None (last day of current year) fillyeargap (bool): Fill year gaps in data. Default: False idblock (str): Block identifier. Default: None idpixeltype (str): Pixel type identifier. Default: None indicatorTypeIds (list): Indicator type IDs — accepted values are 1-5 only (1 VVI, 2/3 weather observed, 4/5 weather forecast). Default: None (auto-set to [1] when index='vegetation-vigor-index') use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): amu_id (required — regional AMU entity ID)
Output columns
entity_id, date, year, regional index values
get_new_token 🔗
Implements token refresh logic for RegionalExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_regional_parameters 🔗
setup_regional_parameters(index='vegetation-vigor-index', start_date='2025-01-01', end_date=None, fillyeargap=False, idblock=None, idpixeltype=None, indicatorTypeIds=None, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure params for regional analytic extraction for an entity id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
index
|
Type of vegetation index (default: "vegetation-vigor-index") |
'vegetation-vigor-index'
|
|
start_date
|
Start date in YYYY-MM-DD format |
'2025-01-01'
|
|
end_date
|
End date in YYYY-MM-DD format (default: last day of current year) |
None
|
|
fillyeargap
|
Whether to fill year gaps in data |
False
|
|
idblock
|
Block ID (267 or 141) |
None
|
|
idpixeltype
|
Pixel type ID - [1] ALL_VEGETATIONS - [2] SUMMER - [3] WINTER - [7] GRASSLANDS - [8] ALL CROPS - [10] CORN - [11] SOYBEAN - [301] BR SOYBEAN - [400] EUR CORN - [401] EUR WINTER WHEAT - [406] EUR BARLEY - [407] EUR RAPESEED - [402] EUR SURGARBEET |
None
|
|
indicatorTypeIds
|
List of indicator type IDs. Only 1-5 are accepted; any other value raises ValueError: - [1] for VVI - [2] for WEATHER OBSERVED ECMWF all blocks - [3] for WEATHER OBSERVED AROME on over France - [4] for WEATHER FORECAST ECMWF all blocks - [5] for WEATHER FORECAST GFS all blocks WEATHER REANALYSIS (10) and WEATHER RAINFALL ESTIMATES HI-RES / CHIRPS (11) are documented by the API but are not in this extractor's accepted set — widen valid_indicatorTypeIds below before passing them. |
None
|
|
partial_frequency
|
Frequency threshold for partial data |
50
|
safe_convert_amu_id
staticmethod
🔗
Safely convert amu_id to integer, handling various input formats.
Handles: - Already an int: returns as-is - String integer: "141" -> 141 - String tuple representation: '(141;"County";...)' -> 141 - Float: 141.0 -> 141
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
amu_id
|
Union[str, int, float]
|
The AMU ID in any format |
required |
Returns:
| Name | Type | Description |
|---|---|---|
int |
int
|
Cleaned integer AMU ID |
Raises:
| Type | Description |
|---|---|
ValueError
|
If amu_id cannot be converted to int |
Examples:
get_regional_ts_by_id 🔗
Query the regional analytic API to get cumulative vegetation vigor data for a given entity id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain the mapped 'amu_id' column (configurable via column_mapping). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Raw JSON response from API. |
get_regional_ts_by_id_safe 🔗
Safe wrapper around get_regional_ts_by_id(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain the mapped 'amu_id' column (configurable via column_mapping). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_regional_json 🔗
Process API time series response to extract daily and average data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
str | dict
|
API response containing time series data |
required |
entity_id
|
str
|
Entity ID for logging purposes |
'unknown'
|
base_year
|
int
|
Base year for converting dayOfYear codes (default 1900) |
1900
|
Returns:
| Name | Type | Description |
|---|---|---|
tuple |
(observed_df, daily_avg_df) - Two DataFrames with formatted data |
process_single_entity_regional 🔗
Process regional extraction for a single region with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict | Series
|
Entity data containing the mapped 'amu_id' column |
required |
params
|
dict
|
regional parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{ "observed_data": DataFrame or None, # Time series observations "daily_avg_data": DataFrame or None, # Daily averages (climatology) "error": dict or None |
|
|
} |
process_entity_regional_bulk_parallel 🔗
process_entity_regional_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='regional', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity regional data using threads + progress bar, with optional fail-safe retry and filter capabilities. Returns TWO normalized dataframes: observed time series and daily averages (climatology).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain the mapped 'amu_id' column. |
required |
params
|
dict
|
Regional parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "regional". |
'regional'
|
use_cache
|
bool
|
If True, use caching for bulk extraction. Default: None (uses instance setting). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains TWO results DataFrames (observed_data and daily_avg_data), errors list, and summary statistics. |