Benchmark and warning extractors🔗
Generated from the source docstrings by mkdocstrings at build time.
See the API reference overview for the other groups.
ChangeIndexExtractor 🔗
Bases: BaseExtractor
Extracts change detection analytics between two vegetation-index images.
Compares a reference image (given per entity) against the nearest prior image in a configurable look-back window to flag significant change. Useful for detecting rapid field changes (harvest, stress events, tillage, etc.).
Documentation: https://docs.earthdaily.com/agro/library/Change_Index/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_ChangeIndex.ipynb
Args (setup_change_index_parameters):
map_type (str): Vegetation index to analyze. Default: 'NDVI'.
Choices: NDVI, EVI, CVI, GNDVI, NDWI.
collections (list[str]): Satellite collections to search. Default: ['Sentinel-2'].
Choices: 'Sentinel-2', 'Landsat'.
max_period_reference (int): Max days back from reference_date to find the reference image. Default: 7.
max_period_previous (int): Max days before the reference image to look for a previous image. Default: 15.
min_period_previous (int): Min days before the reference image to look for a previous image. Default: 5.
same_sensor (bool): Require both images to come from the same sensor. Default: False.
parameter_profile (str): API parameter profile name. Default: 'change_index_v1'.
publish_af (bool): Whether to publish to AF (includes field_id in the API
payload if True). Default: False.
Entity fields (via column_mapping): id, geometry (required); reference_date (required per entity — YYYY-MM-DD reference image date); crop, sowing_date (optional)
Output columns
entity_id, reference_date, status, + any change metrics returned by the API
setup_change_index_parameters 🔗
setup_change_index_parameters(map_type='NDVI', collections=None, max_period_reference=7, max_period_previous=15, min_period_previous=5, same_sensor=False, parameter_profile='change_index_v1', partial_frequency=50, publish_af=False, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for change index extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
map_type
|
str
|
Vegetation index (NDVI, EVI, CVI, GNDVI, NDWI). Default: 'NDVI'. |
'NDVI'
|
collections
|
list[str]
|
Satellite collections. Default: ['Sentinel-2']. |
None
|
max_period_reference
|
int
|
Max days back from reference_date for reference image. Default: 7. |
7
|
max_period_previous
|
int
|
Max days before reference image for previous image. Default: 15. |
15
|
min_period_previous
|
int
|
Min days before reference image for previous image. Default: 5. |
5
|
same_sensor
|
bool
|
Require same sensor for both images. Default: False. |
False
|
parameter_profile
|
str
|
API parameter profile name. Default: 'change_index_v1'. |
'change_index_v1'
|
partial_frequency
|
int
|
How often to save partial results. Default: 50. |
50
|
publish_af
|
bool
|
Whether to publish to AF (includes |
False
|
column_mapping
|
dict
|
Column name overrides (must include 'reference_date'). |
None
|
output_mapping
|
dict
|
Output column renames. |
None
|
exclude_columns
|
list
|
Columns to exclude from output. |
None
|
output_columns
|
list
|
Whitelist of output columns. |
None
|
use_cache
|
bool
|
Whether to use caching. |
None
|
get_change_index_api 🔗
Call the Change Index API for a single entity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict | Series
|
Entity data with 'id', 'geometry', 'reference_date' and optionally 'crop' and 'sowing_date'. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Parsed JSON response from the Change Index API. |
get_change_index_api_safe 🔗
Safe wrapper around get_change_index_api().
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"success": bool, "data": dict|None, "error": str|None, "entity_id": str} |
format_change_index_json 🔗
Normalize a Change Index API response into a single-row DataFrame.
The API returns metrics nested under a data object (SPAEFIndex,
ChangeIndex, CurrentMap/ReferenceMap stats, date_ref, sensor_ref,
date_current, sensor_current, ...). This method flattens those onto
the row so each metric is its own column.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
API response JSON (augmented with 'id' and 'reference_date' by get_change_index_api). |
required |
params
|
dict
|
Override change index parameters. |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Single-row DataFrame with entity_id, reference_date, |
|
|
status, and one column per metric in the nested |
process_single_entity_change_index 🔗
Process change index extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict | Series
|
Entity data with id, geometry, reference_date. |
required |
params
|
dict
|
Override parameters. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_change_index_bulk_extraction_parallel 🔗
process_change_index_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='change_index', generate_report=False, report_options=None, use_cache=None)
Bulk processing of Change Index extraction using threads + progress bar.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id', 'geometry', and 'reference_date'; column names may be remapped via column_mapping). |
required |
params
|
dict
|
Override change index parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' or 'include'. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. |
None
|
skip_export
|
bool
|
If True, skip final export. Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames. Default: "change_index". |
'change_index'
|
generate_report
|
bool
|
Generate HTML report. Default: False. |
False
|
report_options
|
dict
|
Report configuration options. |
None
|
use_cache
|
bool
|
Enable/disable caching for this call. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
results_df, global_errors, total_entities, total_calculations, successful_calculations, failed_calculations, failed_ids |
InSeasonMonitoringExtractor 🔗
Bases: BaseExtractor
Extracts in-season crop monitoring analytics for agricultural entities.
Combines satellite and weather data for real-time crop performance monitoring. Runs against either resolution (LR or MR) with configurable season windows.
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_InSeasonMonitoring.ipynb
Args (setup_inseason_monitoring_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (str): Target year. Default: '2025' data_source (str): Data source - 'LR' or 'MR'. A single source per run (the API takes one dataSource value); run twice to cover both. Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional)
Output columns
entity_id, + in-season monitoring metrics and alerts
get_new_token 🔗
Implements token refresh logic for InSeasonExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
validate_crop
staticmethod
🔗
Validate and normalize the crop type for in-season monitoring requests.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop
|
str
|
The crop name to validate. |
required |
available_crops
|
list
|
List of accepted crop names (expected in uppercase). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the crop is missing, invalid, or not a string. |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The validated crop name, normalized to uppercase. |
setup_inseason_monitoring_parameters 🔗
setup_inseason_monitoring_parameters(season_duration=120, season_start_day=1, season_start_month=4, year='2025', data_source='LR', partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for in-season monitoring extraction
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
season_duration
|
int
|
Crop cycle duration in days |
120
|
season_start_month
|
int
|
Crop cycle start month (1-12) |
4
|
season_start_day
|
int
|
Crop cycle start day (1-31) |
1
|
year
|
int
|
Year to analyze |
'2025'
|
data_source
|
str
|
Imagery type used to build the time series — 'LR' (low resolution) or 'MR' (medium resolution). One source per run; the API takes a single dataSource value. Default: 'LR' |
'LR'
|
partial_frequency
|
int
|
How often to save partial results |
50
|
get_inseason_monitoring_api 🔗
Request In-Season Monitoring for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
enttiy data (id, geometry, crop) |
required |
get_inseason_monitoring_api_safe 🔗
Safe wrapper around get_inseason_monitoring_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_inseason_json 🔗
Unpack In-Season Monitoring JSON into flat rows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
API response JSON with keys 'id' and 'data' |
required |
Returns:
| Type | Description |
|---|---|
|
list of dict: flattened rows |
process_inseason_bulk_extraction_parallel 🔗
process_inseason_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='inseason', generate_report=False, report_options=None, use_cache=None)
Bulk processing of In-Season Monitoring (ISM) requests using threads + progress bar, with optional fail-safe retry and partial save capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id', 'geometry', and 'crop'). |
required |
params
|
dict
|
Override monitoring parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "inseason". |
'inseason'
|
use_cache
|
bool
|
Enable/disable caching for this call. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] } |
save_inseason_reports 🔗
Save extraction In-Season Monitoring extraction report as CSV.
DifferenceExtractor 🔗
Bases: BaseExtractor
Extracts difference maps between two satellite images for agricultural entities.
Calls the Difference Map API to compute pixel-level vegetation index differences between two dates. Returns field-level statistics (max, mean, min) and per-range breakdowns with pixel counts and area.
Documentation: https://docs.earthdaily.com/agro/library/Field%20Level%20Maps/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_difference.ipynb
Args (setup_difference_parameters): product (str): Vegetation index short name (e.g. 'NDVI', 'EVI'). Default: 'NDVI' output_epsg (int): Output coordinate system. Default: 4326 postprocess (str): Processing mode - 'stats', 'links', or 'file'. Default: 'stats' map_format (str): Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None output_path (str): Directory for file downloads. Required when postprocess='file'. directLinks (bool): Request direct download links from API. Default: False partial_frequency (int): How often to save partial results. Default: 50 use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry, image_id_1, image_id_2 (required)
Output columns
Varies by postprocess mode. stats: entity_id, stat_max, stat_mean, stat_min, range_{n}min, rangemax, rangepixels, rangearea, rangecolor_r, rangecolor_g, rangecolor_b links: image_png_link, worldfile_link, thumbnail_link, worldfile, bbox_, map_width, map_height file: status, map_format, file_count, total_size_bytes, saved_files
TODO
- Cache support: cache is disabled by default because results depend on image_id_1 and image_id_2. To enable cache properly, both image IDs need to be serialized to a stable string for cache key lookup.
setup_difference_parameters 🔗
setup_difference_parameters(product='NDVI', output_epsg=4326, postprocess='stats', map_format=None, output_path=None, skip_existing=True, directLinks=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for difference map extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
product
|
str
|
Vegetation index short name. Default: 'NDVI' Supported values: NDVI, EVI, GNDVI, NDRE, CVI, CVIN, LAI Also accepts full names (e.g. 'DIFFERENCE_NDVI') |
'NDVI'
|
output_epsg
|
int
|
Output EPSG code for projection. Default: 4326 |
4326
|
postprocess
|
str
|
Processing mode - 'stats', 'links', or 'file'. Default: 'stats' |
'stats'
|
map_format
|
str
|
Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None |
None
|
output_path
|
str
|
Directory for file downloads. Required when postprocess='file'. |
None
|
skip_existing
|
bool
|
Skip download if file already exists. Default: True |
True
|
directLinks
|
bool
|
Request direct download links. Default: False Note: Automatically set to True when postprocess='links' |
False
|
partial_frequency
|
int
|
How often to save partial results. Default: 50 |
50
|
column_mapping
|
dict
|
Column name overrides. |
None
|
output_mapping
|
dict
|
Output column renames. |
None
|
exclude_columns
|
list
|
Columns to exclude from output. |
None
|
output_columns
|
list
|
Whitelist of output columns. |
None
|
use_cache
|
bool
|
Whether to use caching. |
None
|
get_difference_map_api 🔗
Call the Difference Map API to compute vegetation index difference between two images. Returns raw Response when map_format is set (file download), or parsed JSON otherwise (stats/links).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id', 'geometry', 'image_id_1', and 'image_id_2'. |
required |
Returns:
| Type | Description |
|---|---|
|
requests.Response (file mode) or dict (stats/links mode) |
get_difference_map_api_safe 🔗
Safe wrapper around get_difference_map_api().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id', 'geometry', 'image_id_1', 'image_id_2' |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{"success": bool, "data": dict|None, "error": str|None, "seasonfield_id": str} |
format_difference_stats_json 🔗
Extract field-level statistics and per-range breakdowns from Difference Map API response.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
Parsed API response |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Single-row DataFrame with stats and flattened range columns |
format_difference_links_json 🔗
Extract direct links and metadata from Difference Map API response.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict or Response
|
API response with _links, worldFile, etc. |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Single-row DataFrame with links and spatial metadata |
format_difference_file 🔗
Save difference map file response (PNG, TIFF.ZIP, or SHP.ZIP) to disk.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id' |
required |
response
|
Response
|
Raw response from get_difference_map_api |
required |
output_path
|
str
|
Directory to save files. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Status and file path information |
process_single_entity_difference 🔗
Process difference map extraction for a single entity with retry logic. Routes to stats, links, or file processing based on postprocess parameter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict | Series
|
Entity data with id, geometry, image_id_1, image_id_2 |
required |
params
|
dict
|
Override parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_entity_difference_bulk_parallel 🔗
process_entity_difference_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='difference', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity difference maps using threads + progress bar, with optional fail-safe retry and filter capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain 'id', 'geometry', 'image_id_1', and 'image_id_2'. |
required |
params
|
dict
|
Difference parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export. Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "difference". |
'difference'
|
generate_report
|
bool
|
Generate HTML report. Default: False. |
False
|
report_options
|
dict
|
Report configuration options. |
None
|
use_cache
|
bool
|
Enable/disable caching for this call. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. |