Skip to content

Risk management extractors🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

HistoricalScoreExtractor 🔗

Bases: BaseExtractor

Extracts historical potential/risk score analytics for agricultural entities.

Computes historical performance-based risk assessment scores by comparing vegetation index patterns across multiple years. Supports configurable season windows, historical season comparison, and multiple detail levels.

Documentation: https://docs.earthdaily.com/agro/library/Historical_Potential_Score/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_historical_score.ipynb

Args (setup_historical_score_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 threshold_start (float): Score threshold start. Default: 0.7 year (int): Target year. Default: 2025 historical_seasons (list): Historical season list for comparison. Default: None data_source (str): Data source ('LR', 'MR'). Default: 'LR' detail_level (str): Output detail ('full' or 'summary'). Default: 'full' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional); historical_seasons (optional)

Output columns

entity_id, score, percentile, rank, + historical comparison metrics

get_new_token 🔗

get_new_token()

Implements token refresh logic for historical_scoreExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_historical_score_parameters 🔗

setup_historical_score_parameters(season_duration=120, season_start_day=1, season_start_month=4, threshold_start=0.7, year=2025, historical_seasons=None, data_source='LR', publish_af=False, partial_frequency=50, detail_level='full', column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for historical_score extraction

Parameters:

Name Type Description Default
season_duration int

Crop cycle duration in days

120
season_start_day int

Crop cycle start day (1-31)

1
season_start_month int

Crop cycle start month (1-12)

4
threshold_start float

Threshold for season start detection (0.0-1.0). Default: 0.7

0.7
year int

Current year to analyze

2025
historical_seasons list or None

List of historical years for comparison (e.g., [2024, 2023, 2022]). If None, no historical comparison. Default: None

None
data_source str

Imagery type - 'LR' (low resolution) or 'MR' (medium resolution)

'LR'
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50
detail_level str

Output detail level - 'summary' (only averages) or 'full' (all per-season data). Default: 'full'

'full'

get_historical_score_api 🔗

get_historical_score_api(entity_data: dict)

Request historical_score for an entity

Parameters:

Name Type Description Default
entity_data dict

Entity data containing: - 'id' (str): Entity identifier - 'geometry' (str): WKT geometry - 'historical_seasons' (list or None, optional): List of historical years (e.g., [2024, 2023, 2022]). If not provided, defaults to None (no historical comparison).

required

Returns:

Name Type Description
dict

API response JSON

get_historical_score_api_safe 🔗

get_historical_score_api_safe(entity_data: dict) -> dict

Safe wrapper around get_historical_score_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_historical_score_json 🔗

format_historical_score_json(response_historical_score_json, params=None, detail_level='full')

Normalize historical_score API response into a clean pandas DataFrame.

Parameters:

Name Type Description Default
response_historical_score_json dict

API response JSON with keys 'id' and 'data'.

required
params dict

historical_score parameters (defaults to self.historical_score_params).

None
detail_level str

Output detail level: - 'summary': Returns only summary metrics (one row) - 'full': Returns summary + all per-season data in one row (wide format). Default: 'full'

'full'

Returns:

Type Description

pd.DataFrame: Normalized historical_score data (single row). - 'summary': Columns: entity_id, average_potential_score, olympic_mean_potential_score, standard_deviation, risk_score - 'full': Summary + potential_score_YYYY, season_break_YYYY for each season

process_historical_score_bulk_extraction_parallel 🔗

process_historical_score_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='historical_score', generate_report=False, report_options=None, use_cache=None)

Bulk processing of historical_score requests using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "historical_score".

'historical_score'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.

InseasonScoreExtractor 🔗

Bases: BaseExtractor

Extracts in-season potential/risk score analytics for agricultural entities.

Computes current-season risk assessment using predictive analytics based on vegetation index trends. Provides real-time scoring and comparison with historical baselines.

Documentation: https://docs.earthdaily.com/agro/library/In-season_Potential_Score/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_inseason_score.ipynb

Args (setup_inseason_score_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 threshold_start (float): Score threshold start. Default: 0.7 year (int): Target year. Default: 2025 data_source (str): Data source ('LR', 'MR'). Default: 'LR' detail_level (str): Output detail ('full' or 'summary'). Default: 'full' historical_seasons (list): Explicit prior season years for the baseline. Default: None nb_historical_year (int): Number of historical years used for the baseline. Default: 1 use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional)

Output columns

entity_id, score, status, trend, + current season metrics

get_new_token 🔗

get_new_token()

Implements token refresh logic for InseasonScoreExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_inseason_score_parameters 🔗

setup_inseason_score_parameters(season_duration=120, season_start_day=1, season_start_month=4, nb_historical_year=1, threshold_start=0.7, historical_seasons=None, data_source='LR', publish_af=False, partial_frequency=50, detail_level='full', column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for inseason_score extraction

Parameters:

Name Type Description Default
season_duration int

Crop cycle duration in days

120
season_start_day int

Crop cycle start day (1-31)

1
season_start_month int

Crop cycle start month (1-12)

4
nb_historical_year int

Number of historical years to include (must be >= 1). Default: 1

1
threshold_start float

Threshold for season start detection (0.0-1.0). Default: 0.7

0.7
historical_seasons list or None

List of historical years for comparison (e.g., [2024, 2023, 2022]). If None, no historical comparison. Default: None

None
data_source str

Imagery type - 'LR' (low resolution) or 'MR' (medium resolution)

'LR'
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50
detail_level str

Output detail level - 'summary' (only averages) or 'full' (all per-season data). Default: 'full'

'full'

get_inseason_score_api 🔗

get_inseason_score_api(entity_data: dict)

Request inseason_score for an entity

Parameters:

Name Type Description Default
entity_data dict

Entity data containing: - 'id' (str): Entity identifier - 'geometry' (str): WKT geometry - 'crop' (str): Crop code (e.g., 'CORN', 'WHEAT', 'OTHERS') - 'sowing_date' (str): Sowing date in YYYY-MM-DD or YYYY-MM-DDTHH:MM:SS format - 'end_date' (str, optional): End date in YYYY-MM-DD or YYYY-MM-DDTHH:MM:SS format. If not provided, calculated as sowing_date + season_duration. - 'historical_seasons' (list or None, optional): List of historical years (e.g., [2024, 2023, 2022]). If not provided, defaults to None (no historical comparison).

required

Returns:

Name Type Description
dict

API response JSON

get_inseason_score_api_safe 🔗

get_inseason_score_api_safe(entity_data: dict) -> dict

Safe wrapper around get_inseason_score_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_inseason_score_json 🔗

format_inseason_score_json(response_inseason_score_json, params=None, detail_level='full')

Normalize inseason_score API response into a clean pandas DataFrame.

Parameters:

Name Type Description Default
response_inseason_score_json dict

API response JSON with keys 'id' and 'data'.

required
params dict

inseason_score parameters (defaults to self.inseason_score_params).

None
detail_level str

Output detail level: - 'summary': Returns only the three main scores (one row) - 'full': Returns all scores (same as summary for in-season score). Default: 'full'

'full'

Returns:

Type Description

pd.DataFrame: Normalized inseason_score data (single row). Columns: entity_id, historical_potential_score, inseason_potential_score, relative_potential_score

process_single_entity_inseason_score 🔗

process_single_entity_inseason_score(row, params=None)

Process a single entity for in-season score extraction.

Parameters:

Name Type Description Default
row Series or dict

Entity data containing: - 'id': Entity identifier - 'geometry': WKT geometry - 'crop': Crop code - 'sowing_date': Sowing date in YYYY-MM-DD format - 'end_date' (optional): End date in YYYY-MM-DD format - 'historical_seasons' (optional): List of historical years

required
params dict

Processing parameters

None

Returns:

Name Type Description
dict

{'data': DataFrame or None, 'error': dict or None}

process_inseason_score_bulk_extraction_parallel 🔗

process_inseason_score_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='inseason_score', generate_report=False, report_options=None, use_cache=None)

Bulk processing of inseason_score requests using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "inseason_score".

'inseason_score'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.

ZARCExtractor 🔗

Bases: BaseExtractor

Extracts ZARC (Zoneamento Agricola de Risco Climatico) analytics for agricultural entities.

Brazilian Agricultural Climate Risk Zoning compliance service. Validates crop cycle parameters against ZARC risk zones and computes soil water balance metrics.

Documentation: https://docs.earthdaily.com/agro/library/ZARC/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_ZARC.ipynb

Args (setup_zarc_parameters): crop (str): Crop type code. Default: 'OTHERS' nb_days_sowing_emergence (int): Days between sowing and emergence. Default: 20 soil_type (str): Soil type classification. Default: None cycle (str): Crop cycle duration. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional)

Output columns

entity_id, crop, risk_level, sowing_date, emergence_date, + soil water balance metrics

setup_zarc_parameters 🔗

setup_zarc_parameters(crop: str = 'OTHERS', nb_days_sowing_emergence: int = 20, soil_type: Optional[str] = None, cycle: Optional[str] = None, partial_frequency: int = 50, column_mapping: dict = None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None) -> None

Configure ZARC global parameters that apply to all entities.

Parameters:

Name Type Description Default
crop str

Default crop type (e.g., "SOYBEANS", "CORN")

'OTHERS'
nb_days_sowing_emergence int

Default days between sowing and emergence

20
soil_type Optional[str]

Optional default soil type

None
cycle Optional[str]

Optional default crop cycle

None
partial_frequency int

How often to save partial results during bulk processing

50

get_zarc_api 🔗

get_zarc_api(entity_data: dict) -> Dict[str, Any]

Call ZARC API for a specific entity.

Parameters:

Name Type Description Default
entity_data dict

Must contain 'id', 'geometry', 'emergence_date'. Can optionally override 'crop', 'soil_type', 'cycle', 'nb_days_sowing_emergence'.

required

Returns:

Name Type Description
dict Dict[str, Any]

Raw JSON response from API.

format_zarc_json staticmethod 🔗

format_zarc_json(response_json: Dict[str, Any]) -> pd.DataFrame

Format ZARC API response into a pandas DataFrame.

get_zarc_api_safe 🔗

get_zarc_api_safe(entity_data: Dict[str, Any]) -> Dict[str, Any]

Safe wrapper around get_zarc_api that returns a structured dict instead of raising exceptions.

Parameters:

Name Type Description Default
entity_data dict

Entity data containing 'id', 'geometry', 'emergence_date', etc.

required

Returns:

Name Type Description
dict Dict[str, Any]

{"success": bool, "data": dict or None, "error": str or None, "entity_id": str}

process_single_entity_zarc 🔗

process_single_entity_zarc(row, params: Optional[Dict[str, Any]] = None) -> Dict[str, Any]

Process ZARC for a single entity with retry and normalization. Uses per-row values when present, otherwise falls back to values passed via params.

Parameters:

Name Type Description Default
row dict | Series

Entity data containing id, geometry, emergence_date

required
params dict

Fallback parameters for nb_days_sowing_emergence

None

Returns:

Name Type Description
dict Dict[str, Any]

{"data": DataFrame or None, "error": dict or None}

process_zarc_bulk_extraction_parallel 🔗

process_zarc_bulk_extraction_parallel(entity_list: DataFrame, max_workers: int = 5, output_path: Optional[str] = None, partial_frequency: int = 50, fail_safe: bool = False, filter_column: Optional[str] = None, filter_value: Optional[Any] = None, filter_type: str = 'exclude', merge_existing: Optional[str] = None, skip_export: bool = False, prefix: str = 'zarc', generate_report: bool = False, report_options: Optional[Dict[str, Any]] = None, use_cache=None) -> Dict[str, Any]

Bulk processing for ZARC with parallel execution, filtering, and export capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

DataFrame with entities to process (must contain 'id', 'geometry', 'emergence_date')

required
max_workers int

Number of parallel threads

5
output_path Optional[str]

Directory to save final CSV results

None
partial_frequency int

How often to save partial results

50
fail_safe bool

If True, retries only previously failed entities

False
filter_column Optional[str]

Column name to filter entities by

None
filter_value Optional[Any]

Value to filter in the filter_column

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only those

'exclude'
merge_existing Optional[str]

Merge strategy - 'auto', 'preserve', or 'mark'

None
skip_export bool

If True, skip final export (useful when chaining extractions)

False
prefix str

Prefix for output filenames and failed IDs files

'zarc'
use_cache

If True, use caching for bulk extraction. Default: None (uses instance setting).

None

Returns:

Name Type Description
dict Dict[str, Any]

Contains results DataFrame, errors list, and summary statistics