Skip to content

Benchmark and warning extractors🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

ChangeIndexExtractor 🔗

Bases: BaseExtractor

Extracts change detection analytics between two vegetation-index images.

Compares a reference image (given per entity) against the nearest prior image in a configurable look-back window to flag significant change. Useful for detecting rapid field changes (harvest, stress events, tillage, etc.).

Documentation: https://docs.earthdaily.com/agro/library/Change_Index/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_ChangeIndex.ipynb

Args (setup_change_index_parameters): map_type (str): Vegetation index to analyze. Default: 'NDVI'. Choices: NDVI, EVI, CVI, GNDVI, NDWI. collections (list[str]): Satellite collections to search. Default: ['Sentinel-2']. Choices: 'Sentinel-2', 'Landsat'. max_period_reference (int): Max days back from reference_date to find the reference image. Default: 7. max_period_previous (int): Max days before the reference image to look for a previous image. Default: 15. min_period_previous (int): Min days before the reference image to look for a previous image. Default: 5. same_sensor (bool): Require both images to come from the same sensor. Default: False. parameter_profile (str): API parameter profile name. Default: 'change_index_v1'. publish_af (bool): Whether to publish to AF (includes field_id in the API payload if True). Default: False.

Entity fields (via column_mapping): id, geometry (required); reference_date (required per entity — YYYY-MM-DD reference image date); crop, sowing_date (optional)

Output columns

entity_id, reference_date, status, + any change metrics returned by the API

get_new_token 🔗

get_new_token()

Implements token refresh logic for ChangeIndexExtractor.

setup_change_index_parameters 🔗

setup_change_index_parameters(map_type='NDVI', collections=None, max_period_reference=7, max_period_previous=15, min_period_previous=5, same_sensor=False, parameter_profile='change_index_v1', partial_frequency=50, publish_af=False, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for change index extraction.

Parameters:

Name Type Description Default
map_type str

Vegetation index (NDVI, EVI, CVI, GNDVI, NDWI). Default: 'NDVI'.

'NDVI'
collections list[str]

Satellite collections. Default: ['Sentinel-2'].

None
max_period_reference int

Max days back from reference_date for reference image. Default: 7.

7
max_period_previous int

Max days before reference image for previous image. Default: 15.

15
min_period_previous int

Min days before reference image for previous image. Default: 5.

5
same_sensor bool

Require same sensor for both images. Default: False.

False
parameter_profile str

API parameter profile name. Default: 'change_index_v1'.

'change_index_v1'
partial_frequency int

How often to save partial results. Default: 50.

50
publish_af bool

Whether to publish to AF (includes field_id in payload if True). Default: False. Mirrors the flag used by the rest of the processor family (Baresoil, Covercrop, Emergence, Greenness, Harvest, Planted, score, InSeasonMonitoring).

False
column_mapping dict

Column name overrides (must include 'reference_date').

None
output_mapping dict

Output column renames.

None
exclude_columns list

Columns to exclude from output.

None
output_columns list

Whitelist of output columns.

None
use_cache bool

Whether to use caching.

None

get_change_index_api 🔗

get_change_index_api(entity_data)

Call the Change Index API for a single entity.

Parameters:

Name Type Description Default
entity_data dict | Series

Entity data with 'id', 'geometry', 'reference_date' and optionally 'crop' and 'sowing_date'.

required

Returns:

Name Type Description
dict

Parsed JSON response from the Change Index API.

get_change_index_api_safe 🔗

get_change_index_api_safe(entity_data)

Safe wrapper around get_change_index_api().

Returns:

Name Type Description
dict

{"success": bool, "data": dict|None, "error": str|None, "entity_id": str}

format_change_index_json 🔗

format_change_index_json(response_json, params=None)

Normalize a Change Index API response into a single-row DataFrame.

The API returns metrics nested under a data object (SPAEFIndex, ChangeIndex, CurrentMap/ReferenceMap stats, date_ref, sensor_ref, date_current, sensor_current, ...). This method flattens those onto the row so each metric is its own column.

Parameters:

Name Type Description Default
response_json dict

API response JSON (augmented with 'id' and 'reference_date' by get_change_index_api).

required
params dict

Override change index parameters.

None

Returns:

Type Description

pd.DataFrame: Single-row DataFrame with entity_id, reference_date,

status, and one column per metric in the nested data object.

process_single_entity_change_index 🔗

process_single_entity_change_index(row, params=None)

Process change index extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict | Series

Entity data with id, geometry, reference_date.

required
params dict

Override parameters.

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_change_index_bulk_extraction_parallel 🔗

process_change_index_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='change_index', generate_report=False, report_options=None, use_cache=None)

Bulk processing of Change Index extraction using threads + progress bar.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'reference_date'; column names may be remapped via column_mapping).

required
params dict

Override change index parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' or 'include'. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'.

None
skip_export bool

If True, skip final export. Default: False.

False
prefix str

Prefix for output filenames. Default: "change_index".

'change_index'
generate_report bool

Generate HTML report. Default: False.

False
report_options dict

Report configuration options.

None
use_cache bool

Enable/disable caching for this call.

None

Returns:

Name Type Description
dict

results_df, global_errors, total_entities, total_calculations, successful_calculations, failed_calculations, failed_ids

InSeasonMonitoringExtractor 🔗

Bases: BaseExtractor

Extracts in-season crop monitoring analytics for agricultural entities.

Combines satellite and weather data for real-time crop performance monitoring. Runs against either resolution (LR or MR) with configurable season windows.

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_InSeasonMonitoring.ipynb

Args (setup_inseason_monitoring_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (str): Target year. Default: '2025' data_source (str): Data source - 'LR' or 'MR'. A single source per run (the API takes one dataSource value); run twice to cover both. Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional)

Output columns

entity_id, + in-season monitoring metrics and alerts

get_new_token 🔗

get_new_token()

Implements token refresh logic for InSeasonExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

validate_crop staticmethod 🔗

validate_crop(crop: str, available_crops: set)

Validate and normalize the crop type for in-season monitoring requests.

Parameters:

Name Type Description Default
crop str

The crop name to validate.

required
available_crops list

List of accepted crop names (expected in uppercase).

required

Raises:

Type Description
ValueError

If the crop is missing, invalid, or not a string.

Returns:

Name Type Description
str

The validated crop name, normalized to uppercase.

setup_inseason_monitoring_parameters 🔗

setup_inseason_monitoring_parameters(season_duration=120, season_start_day=1, season_start_month=4, year='2025', data_source='LR', partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for in-season monitoring extraction

Parameters:

Name Type Description Default
season_duration int

Crop cycle duration in days

120
season_start_month int

Crop cycle start month (1-12)

4
season_start_day int

Crop cycle start day (1-31)

1
year int

Year to analyze

'2025'
data_source str

Imagery type used to build the time series — 'LR' (low resolution) or 'MR' (medium resolution). One source per run; the API takes a single dataSource value. Default: 'LR'

'LR'
partial_frequency int

How often to save partial results

50

get_inseason_monitoring_api 🔗

get_inseason_monitoring_api(entity_data: dict)

Request In-Season Monitoring for an entity

Parameters:

Name Type Description Default
entity_data dict

enttiy data (id, geometry, crop)

required

get_inseason_monitoring_api_safe 🔗

get_inseason_monitoring_api_safe(entity_data: dict) -> dict

Safe wrapper around get_inseason_monitoring_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_inseason_json 🔗

format_inseason_json(response_monitoring_json)

Unpack In-Season Monitoring JSON into flat rows.

Parameters:

Name Type Description Default
response_json dict

API response JSON with keys 'id' and 'data'

required

Returns:

Type Description

list of dict: flattened rows

process_inseason_bulk_extraction_parallel 🔗

process_inseason_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='inseason', generate_report=False, report_options=None, use_cache=None)

Bulk processing of In-Season Monitoring (ISM) requests using threads + progress bar, with optional fail-safe retry and partial save capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "inseason".

'inseason'
use_cache bool

Enable/disable caching for this call.

None

Returns:

Name Type Description
dict

{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] }

save_inseason_reports 🔗

save_inseason_reports(results_df, global_errors, output_path)

Save extraction In-Season Monitoring extraction report as CSV.

DifferenceExtractor 🔗

Bases: BaseExtractor

Extracts difference maps between two satellite images for agricultural entities.

Calls the Difference Map API to compute pixel-level vegetation index differences between two dates. Returns field-level statistics (max, mean, min) and per-range breakdowns with pixel counts and area.

Documentation: https://docs.earthdaily.com/agro/library/Field%20Level%20Maps/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_difference.ipynb

Args (setup_difference_parameters): product (str): Vegetation index short name (e.g. 'NDVI', 'EVI'). Default: 'NDVI' output_epsg (int): Output coordinate system. Default: 4326 postprocess (str): Processing mode - 'stats', 'links', or 'file'. Default: 'stats' map_format (str): Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None output_path (str): Directory for file downloads. Required when postprocess='file'. directLinks (bool): Request direct download links from API. Default: False partial_frequency (int): How often to save partial results. Default: 50 use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry, image_id_1, image_id_2 (required)

Output columns

Varies by postprocess mode. stats: entity_id, stat_max, stat_mean, stat_min, range_{n}min, rangemax, rangepixels, rangearea, rangecolor_r, rangecolor_g, rangecolor_b links: image_png_link, worldfile_link, thumbnail_link, worldfile, bbox_, map_width, map_height file: status, map_format, file_count, total_size_bytes, saved_files

TODO
  • Cache support: cache is disabled by default because results depend on image_id_1 and image_id_2. To enable cache properly, both image IDs need to be serialized to a stable string for cache key lookup.

get_new_token 🔗

get_new_token()

Implements token refresh logic for DifferenceExtractor.

setup_difference_parameters 🔗

setup_difference_parameters(product='NDVI', output_epsg=4326, postprocess='stats', map_format=None, output_path=None, skip_existing=True, directLinks=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for difference map extraction.

Parameters:

Name Type Description Default
product str

Vegetation index short name. Default: 'NDVI' Supported values: NDVI, EVI, GNDVI, NDRE, CVI, CVIN, LAI Also accepts full names (e.g. 'DIFFERENCE_NDVI')

'NDVI'
output_epsg int

Output EPSG code for projection. Default: 4326

4326
postprocess str

Processing mode - 'stats', 'links', or 'file'. Default: 'stats'

'stats'
map_format str

Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None

None
output_path str

Directory for file downloads. Required when postprocess='file'.

None
skip_existing bool

Skip download if file already exists. Default: True

True
directLinks bool

Request direct download links. Default: False Note: Automatically set to True when postprocess='links'

False
partial_frequency int

How often to save partial results. Default: 50

50
column_mapping dict

Column name overrides.

None
output_mapping dict

Output column renames.

None
exclude_columns list

Columns to exclude from output.

None
output_columns list

Whitelist of output columns.

None
use_cache bool

Whether to use caching.

None

get_difference_map_api 🔗

get_difference_map_api(entity_data: dict)

Call the Difference Map API to compute vegetation index difference between two images. Returns raw Response when map_format is set (file download), or parsed JSON otherwise (stats/links).

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id', 'geometry', 'image_id_1', and 'image_id_2'.

required

Returns:

Type Description

requests.Response (file mode) or dict (stats/links mode)

get_difference_map_api_safe 🔗

get_difference_map_api_safe(entity_data: dict) -> dict

Safe wrapper around get_difference_map_api().

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id', 'geometry', 'image_id_1', 'image_id_2'

required

Returns:

Name Type Description
dict dict

{"success": bool, "data": dict|None, "error": str|None, "seasonfield_id": str}

format_difference_stats_json 🔗

format_difference_stats_json(response_json)

Extract field-level statistics and per-range breakdowns from Difference Map API response.

Parameters:

Name Type Description Default
response_json dict

Parsed API response

required

Returns:

Type Description

pd.DataFrame: Single-row DataFrame with stats and flattened range columns

format_difference_links_json(response_json)

Extract direct links and metadata from Difference Map API response.

Parameters:

Name Type Description Default
response_json dict or Response

API response with _links, worldFile, etc.

required

Returns:

Type Description

pd.DataFrame: Single-row DataFrame with links and spatial metadata

format_difference_file 🔗

format_difference_file(entity_data: dict, response, output_path=None)

Save difference map file response (PNG, TIFF.ZIP, or SHP.ZIP) to disk.

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id'

required
response Response

Raw response from get_difference_map_api

required
output_path str

Directory to save files.

None

Returns:

Name Type Description
dict

Status and file path information

process_single_entity_difference 🔗

process_single_entity_difference(row, params=None)

Process difference map extraction for a single entity with retry logic. Routes to stats, links, or file processing based on postprocess parameter.

Parameters:

Name Type Description Default
row dict | Series

Entity data with id, geometry, image_id_1, image_id_2

required
params dict

Override parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_entity_difference_bulk_parallel 🔗

process_entity_difference_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='difference', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity difference maps using threads + progress bar, with optional fail-safe retry and filter capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain 'id', 'geometry', 'image_id_1', and 'image_id_2'.

required
params dict

Difference parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export. Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "difference".

'difference'
generate_report bool

Generate HTML report. Default: False.

False
report_options dict

Report configuration options.

None
use_cache bool

Enable/disable caching for this call.

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics.