Skip to content

Crop development and stressors extractors🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

DiseaseExtractor 🔗

Bases: BaseExtractor

Extracts crop disease risk analytics for agricultural entities.

Evaluates disease risk based on weather conditions and crop type. Returns daily disease parameter values (e.g., infection risk, sporulation, severity) for supported crops (corn, soybeans).

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_disease.ipynb

Args (setup_disease_parameters): start_date (str): Start date in YYYY-MM-DD format. Default: None end_date (str): End date in YYYY-MM-DD format. Default: None kpi_filter (dict): KPI aggregation rules applied to the daily disease series. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop, start_date, end_date (optional overrides)

Output columns

entity_id, date, + dynamic disease parameter columns (flattened from API response)

get_new_token 🔗

get_new_token()

Implements token refresh logic for DiseaseExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_disease_parameters 🔗

setup_disease_parameters(start_date=None, end_date=None, partial_frequency=50, kpi_filter=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure params for disease extraction for an entity

Parameters:

Name Type Description Default
start_date str

Default start date (YYYY-MM-DD) used when entity row has no start_date

None
end_date str

Default end date (YYYY-MM-DD) used when entity row has no end_date

None
partial_frequency int

How often to save partial results

50
kpi_filter dict

KPI aggregation config. See api_utils.validate_kpi_filter for the supported aggregations and their threshold/window contracts.

None
column_mapping dict

Override input column mapping for this extractor. Example: {"id": "entity_id", "geometry": "wkt", "crop": "crop_type"}

None

get_disease_data 🔗

get_disease_data(entity_data)

Get disease data for selected parameters using POST request

Parameters:

Name Type Description Default
entity_data dict

Entity data with id, geometry, start_date, end_date

required

Returns:

Name Type Description
dict

Disease API response

get_disease_data_safe 🔗

get_disease_data_safe(entity_data: dict) -> dict

Safe wrapper around get_disease_data(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain 'id', 'geometry', 'start_date', 'end_date' keys

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str

dict

}

format_disease_json 🔗

format_disease_json(response_disease_json, entity_data: dict = None)

Format disease API response into a DataFrame. Dynamically handles all disease parameters returned by the API.

Parameters:

Name Type Description Default
response_disease_json list or dict

JSON response from get_disease_data

required
entity_data dict

Original entity data to include ID in output

None

Returns:

Type Description

pd.DataFrame: DataFrame with date and all disease parameter columns

process_single_entity_disease 🔗

process_single_entity_disease(row, params=None)

Process a single entity for disease data extraction with optional KPI filtering.

Parameters:

Name Type Description Default
row pd.DataFrame, pd.Series, or dict

Entity data with required fields

required
params dict

Override disease_params

None

Returns:

Name Type Description
dict

Contains 'data' (DataFrame) and 'error'

process_entity_disease_bulk_parallel 🔗

process_entity_disease_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='disease', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity disease data using threads + progress bar, with optional fail-safe retry and filter capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process

required
params dict

Disease parameters override

None
max_workers int

Number of threads to use

5
output_path str

Directory to save results

None
partial_frequency int

How often to save partial results

50
fail_safe bool

If True, retries only previously failed entities

False
filter_column str

Column name to filter entities by

None
filter_value any

Value to filter

None
filter_type str

'exclude' or 'include'

'exclude'
merge_existing str

Merge strategy

None
skip_export bool

If True, skip final export

False
prefix str

Prefix for output filenames

'disease'
use_cache bool

Enable/disable caching for this call

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics

EmergenceExtractor 🔗

Bases: BaseExtractor

Detects crop emergence dates from imagery derived vegetation index time series.

For each field (id + geometry + crop) returns an estimated emergence date with a confidence score and status. Supports INSEASON, HISTORICAL and DELAY detection modes and LR / MR data sources. The season window comes from setup (season_*, year), never from a per-entity sowing date.

Documentation: https://docs.earthdaily.com/agro/library/Emergence/ Notebook: EDAgriculture_Emergence_Processor_Function_Dev.ipynb API endpoint: {agro_urls["emergence_url"][env]}/... (env-dependent)

Args (setup_emergence_parameters): emergence_type (str): Detection type ('INSEASON', 'HISTORICAL', 'DELAY'). Default: 'INSEASON' season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 data_source (str): Data source ('LR', 'MR'). Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

For HISTORICAL, a per-field historical_seasons input (via column_mapping) lists the calendar years a field actually grew the crop, as a comma-separated string ("2020,2022,2024") or a list of ints. When supplied, emergence_year_N is kept only for those years (others nulled) and a separate avg_emergence_matching_years column holds the average recomputed over the retained years (MM-DD). historical_average_emergence (the raw API average over all 5 seasons) is left untouched.

Entity fields (via column_mapping): id, geometry, crop (required); historical_seasons (optional — HISTORICAL only). sowing_date is NOT read: the season window comes from setup.

Output columns

entity_id, emergence_date, confidence, status, emergence_year_1..5, historical_average_emergence, avg_emergence_matching_years (MM-DD)

get_new_token 🔗

get_new_token()

Implements token refresh logic for EmergenceExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

validate_crop staticmethod 🔗

validate_crop(crop: str, available_crops: set)

Validate and normalize the crop type for emergence requests.

Parameters:

Name Type Description Default
crop str

The crop name to validate.

required
available_crops list

List of accepted crop names (expected in uppercase).

required

Raises:

Type Description
ValueError

If the crop is missing, invalid, or not a string.

Returns:

Name Type Description
str

The validated crop name, normalized to uppercase.

setup_emergence_parameters 🔗

setup_emergence_parameters(emergence_type='INSEASON', season_duration=120, season_start_day=1, season_start_month=4, year=2025, data_source='LR', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for Emergence extraction

Parameters:

Name Type Description Default
emergence_type str

Selection between INSEASON, HISTORICAL or DELAY

'INSEASON'
season_duration int

Crop cycle duration in days

120
season_start_day int

Crop cycle start day (1-31)

1
season_start_month int

Crop cycle start month (1-12)

4
year int

Year to analyze

2025
data_source str

Imagery type - 'LR' (low resolution) or 'MR' (medium resolution)

'LR'
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50

get_emergence_api 🔗

get_emergence_api(entity_data: dict)

Request Emergence for an entity

Parameters:

Name Type Description Default
entity_data dict

entity data (id, geometry, crop)

required

get_emergence_api_safe 🔗

get_emergence_api_safe(entity_data: dict) -> dict

Safe wrapper around get_emergence_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_emergence_json 🔗

format_emergence_json(response_emergence_json, params=None, historical_seasons=None)

Normalize Emergence API response into a clean pandas DataFrame. Supports multiple emergence types: - INSEASON - HISTORICAL - DELAY

Parameters:

Name Type Description Default
response_emergence_json dict

API response JSON with keys 'id' and 'data'.

required
params dict

Emergence parameters (defaults to self.emergence_params).

None
historical_seasons str | list | None

HISTORICAL only — the calendar years a field actually grew the target crop, as a comma-separated string ("2020,2022,2024") or a list of ints. When supplied, emergence_year_N is kept only for the retained years (others set to None) and a separate avg_emergence_matching_years column holds the average recomputed over the retained years. When None (default), no filtering is applied and the new column is null. historical_average_emergence is left untouched in all cases.

None

Returns:

Type Description

pd.DataFrame: Flat normalized emergence data. Columns depend on emergence_type.

process_single_entity_emergence 🔗

process_single_entity_emergence(row, params=None)

Process emergence extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict

Entity data containing id, geometry, and crop

required
params dict

Emergence parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_emergence_bulk_extraction_parallel 🔗

process_emergence_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='emergence', generate_report=False, report_options=None, use_cache=None)

Bulk processing of emergence requests using threads + progress bar, with optional fail-safe retry, partial save capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "emergence".

'emergence'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] }

GreennessExtractor 🔗

Bases: BaseExtractor

Extracts crop greenness (vegetation vigor) analytics for agricultural entities.

Monitors crop canopy development and vegetation vigor through greenness detection. Evaluates crop health status based on satellite-derived vegetation indices.

Documentation: https://docs.earthdaily.com/agro/library/greenness/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_greenness.ipynb

Args (setup_greenness_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 sowing_date (str): Sowing date in YYYY-MM-DD format. Default: '2025-04-01' data_source (str): Data source ('LR', 'MR'). Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional)

Output columns

entity_id, greenness_date, confidence, status, + detection metadata

get_new_token 🔗

get_new_token()

Implements token refresh logic for GreennessExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

validate_crop staticmethod 🔗

validate_crop(crop: str, available_crops: set)

Validate and normalize the crop type for greenness requests.

Parameters:

Name Type Description Default
crop str

The crop name to validate.

required
available_crops set

Set of accepted crop names (expected in uppercase).

required

Raises:

Type Description
ValueError

If the crop is missing, invalid, or not a string.

Returns:

Name Type Description
str

The validated crop name, normalized to uppercase.

setup_greenness_parameters 🔗

setup_greenness_parameters(season_duration=120, season_start_day=1, season_start_month=4, year=2025, sowing_date='2025-04-01', data_source='LR', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for greenness detection extraction.

Parameters:

Name Type Description Default
season_duration int

Crop cycle duration in days

120
season_start_day int

Crop cycle start day (1-31)

1
season_start_month int

Crop cycle start month (1-12)

4
year int

Year to analyze

2025
sowing_date str

Sowing date in YYYY-MM-DD format

'2025-04-01'
data_source str

Imagery type - 'LR' (low resolution) or 'MR' (medium resolution)

'LR'
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50

get_greenness_api 🔗

get_greenness_api(entity_data: dict)

Request greenness detection for an entity.

Parameters:

Name Type Description Default
entity_data dict

entity data (id, geometry, crop)

required

get_greenness_api_safe 🔗

get_greenness_api_safe(entity_data: dict) -> dict

Safe wrapper around get_greenness_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_greenness_json 🔗

format_greenness_json(response_greenness_json, params=None)

Normalize Greenness Detection API response into a clean pandas DataFrame.

Parameters:

Name Type Description Default
response_greenness_json dict

API response JSON.

required
params dict

Greenness parameters (defaults to self.greenness_params).

None

Returns:

Type Description

pd.DataFrame: Flat normalized greenness data.

process_single_entity_greenness 🔗

process_single_entity_greenness(row, params=None)

Process greenness extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict

Entity data containing id, geometry, and crop

required
params dict

Greenness parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_greenness_bulk_extraction_parallel 🔗

process_greenness_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='greenness', generate_report=False, report_options=None, use_cache=None)

Bulk processing of greenness requests using threads + progress bar, with optional fail-safe retry, partial save capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override greenness parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "greenness".

'greenness'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] }

HarvestExtractor 🔗

Bases: BaseExtractor

Extracts harvest detection analytics for agricultural entities.

Identifies harvest dates and monitors harvest progress using satellite-based temporal analysis. Supports in-season, historical and harvest-readiness detection.

Documentation: https://docs.earthdaily.com/agro/library/Harvest_Detection/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_harvest.ipynb

Args (setup_harvest_parameters): harvest_type (str): Detection type - 'INSEASON_HARVEST', 'HISTORICAL_HARVEST' or 'HARVEST_READINESS'. Default: 'INSEASON_HARVEST' season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 data_source (str): Data source ('LR', 'MR'). Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

For HISTORICAL_HARVEST, a per-field historical_seasons input (via column_mapping) lists the calendar years a field actually grew the crop, as a comma-separated string ("2020,2022,2024") or a list of ints. When supplied, harvest_year_N is kept only for those years (others nulled) and a separate avg_harvest_matching_years column holds the average recomputed over the retained years (MM-DD). historical_harvest_average (the raw API average over all 5 seasons) is left untouched.

Entity fields (via column_mapping): id, geometry, crop (required); historical_seasons (optional — HISTORICAL_HARVEST only). sowing_date is NOT read: the season window comes from setup.

Output columns

Varies by harvest_type. INSEASON_HARVEST: entity_id, harvest_date, harvest_status HISTORICAL_HARVEST: entity_id, harvest_year_1..5, historical_harvest_average, avg_harvest_matching_years (MM-DD) HARVEST_READINESS: entity_id, harvest_readiness_date, is_ready

get_new_token 🔗

get_new_token()

Implements token refresh logic for HarvestExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

validate_crop staticmethod 🔗

validate_crop(crop: str, available_crops: set)

Validate and normalize the crop type for harvest requests.

Parameters:

Name Type Description Default
crop str

The crop name to validate.

required
available_crops list

List of accepted crop names (expected in uppercase).

required

Raises:

Type Description
ValueError

If the crop is missing, invalid, or not a string.

Returns:

Name Type Description
str

The validated crop name, normalized to uppercase.

setup_harvest_parameters 🔗

setup_harvest_parameters(harvest_type='INSEASON_HARVEST', season_duration=120, season_start_day=1, season_start_month=4, year=2025, data_source='LR', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for harvest extraction

Parameters:

Name Type Description Default
harvest_type str

One of 'INSEASON_HARVEST' (harvest date + status for the current season), 'HISTORICAL_HARVEST' (harvest date for each of the last 5 seasons) or 'HARVEST_READINESS' (readiness date + is_ready flag)

'INSEASON_HARVEST'
season_duration int

Crop cycle duration in days

120
season_start_day int

Crop cycle start day (1-31)

1
season_start_month int

Crop cycle start month (1-12)

4
year int

Year to analyze

2025
data_source str

Imagery type - 'LR' (low resolution) or 'MR' (medium resolution)

'LR'
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50

get_harvest_api 🔗

get_harvest_api(entity_data: dict)

Request harvest for an entity

Parameters:

Name Type Description Default
entity_data dict

entity data (id, geometry, crop)

required

get_harvest_api_safe 🔗

get_harvest_api_safe(entity_data: dict) -> dict

Safe wrapper around get_harvest_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_harvest_json 🔗

format_harvest_json(response_harvest_json, params=None, historical_seasons=None)

Normalize Harvest API response into a clean pandas DataFrame. Supports multiple harvest types: - INSEASON_HARVEST - HISTORICAL_HARVEST - HARVEST_READINESS

Parameters:

Name Type Description Default
response_harvest_json dict

API response JSON with keys 'id' and 'data'.

required
params dict

Harvest parameters (defaults to self.harvest_params).

None
historical_seasons str | list | None

HISTORICAL_HARVEST only — the calendar years a field actually grew the target crop, as a comma-separated string ("2020,2022,2024") or a list of ints. When supplied, harvest_year_N is kept only for the retained years (others set to None) and a separate avg_harvest_matching_years column holds the average recomputed over the retained years (MM-DD). When None (default), no filtering is applied and the new column is null. historical_harvest_average is left untouched in all cases.

None

Returns:

Type Description

pd.DataFrame: Flat normalized harvest data. Columns depend on harvest_type.

process_single_entity_harvest 🔗

process_single_entity_harvest(row, params=None)

Process harvest extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict

Entity data containing id, geometry, and crop

required
params dict

Harvest parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_harvest_bulk_extraction_parallel 🔗

process_harvest_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='harvest', generate_report=False, report_options=None, use_cache=None)

Bulk processing of harvest requests using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "harvest".

'harvest'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.

PlantedExtractor 🔗

Bases: BaseExtractor

Extracts planted area estimation analytics for agricultural entities.

Estimates planted acreage using crop identification and emergence data. Validates planting status and computes planted area percentages.

Documentation: https://docs.earthdaily.com/agro/library/Planted_Area/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_coverage.ipynb

Args (setup_planted_parameters): processor_mode (str): Processing mode - 'PLANTED_AREA' (planted area and percentage) or 'CONTROL' (compares the field against control_threshold). Default: 'PLANTED_AREA' emergence_date (str): Default emergence date in YYYY-MM-DD. Default: None (entities may override via the emergence_date row column) threshold (int): Decision threshold in days. Default: 120 control_threshold (int | float): Percentage threshold, CONTROL mode only. Sent to the API divided by 100, so the default 4 becomes controlThreshold=0.04. Default: 4 use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); emergence_date (optional per-entity override of params.emergence_date — wire the EmergenceExtractor output in via column_mapping); crop, sowing_date (optional)

Output columns

Varies by processor_mode. PLANTED_AREA: entity_id, planted_area_m2, planted_percentage CONTROL: entity_id, difference, control_threshold, control_result

get_new_token 🔗

get_new_token()

Implements token refresh logic for PlantedExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_planted_parameters 🔗

setup_planted_parameters(processor_mode='PLANTED_AREA', emergence_date=None, threshold=120, control_threshold=4, publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for planted area extraction

Parameters:

Name Type Description Default
processor_mode str

Selection between 'PLANTED_AREA' or 'CONTROL'

'PLANTED_AREA'
emergence_date str

Default emergence date in YYYY-MM-DD. May be overridden per entity via the emergence_date row column (column_mapping-aware, so upstream Emergence output can be wired in). If omitted here, every row must provide its own.

None
threshold int

Days between images before emergence and second image (default: 120)

120
control_threshold float

Percentage threshold for CONTROL mode (default: 4, used as 0.04)

4
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50

get_planted_api 🔗

get_planted_api(entity_data: dict)

Request planted area calculation for an entity

Parameters:

Name Type Description Default
entity_data dict

entity data (id, geometry)

required

get_planted_api_safe 🔗

get_planted_api_safe(entity_data: dict) -> dict

Safe wrapper around get_planted_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str

dict

}

format_planted_json 🔗

format_planted_json(response_planted_json, params=None)

Normalize Planted Area API response into a clean pandas DataFrame. Supports two processor modes: - PLANTED_AREA: Returns planted area in m² and percentage - CONTROL: Returns difference, control threshold, and result boolean

Parameters:

Name Type Description Default
response_planted_json dict

API response JSON.

required
params dict

Planted area parameters (defaults to self.planted_params).

None

Returns:

Type Description

pd.DataFrame: Flat normalized planted area data. Columns depend on processor_mode.

process_single_entity_planted 🔗

process_single_entity_planted(row, params=None)

Process planted area extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict

Entity data containing id and geometry

required
params dict

Planted area parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_planted_bulk_extraction_parallel 🔗

process_planted_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='planted', generate_report=False, report_options=None, use_cache=None)

Bulk processing of planted area requests using threads + progress bar, with optional fail-safe retry and partial save capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id' and 'geometry').

required
params dict

Override planted area parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "planted".

'planted'
use_cache bool

Enable/disable caching for this call.

None

Returns:

Name Type Description
dict

{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] }

GDDExtractor 🔗

Bases: BaseExtractor

Extracts GDD (Growing Degree Days) analytics for agricultural entities.

Computes daily and cumulated Growing Degree Days from a start date to a last date, using a lower (and optional upper) temperature threshold. The output is a long DataFrame: one row per entity per day, with columns for minimum/maximum temperature, daily GDD and cumulated GDD.

Endpoint: {weather_url}/analytics/gdd

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgro_growing_degree_days.ipynb

Args (setup_gdd_parameters): provider (str): Weather data provider. Default: 'GLOBAL1' lower_threshold (float): Lower GDD temperature threshold (required by API). Default: 10.0 upper_threshold (float): Optional upper temperature threshold (capped GDD). Default: None start_date (str): Period start in YYYY-MM-DD (maps to API StartDate). Default: None end_date (str): Period end in YYYY-MM-DD (maps to API LastDate). Default: None reset_cumulative_every_year (bool): Reset the cumulated GDD at each year boundary. Default: False extrapolate_forecast_data (bool): Extend the series with forecast data. Default: False use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required) start_date → optional per-entity override for the API StartDate end_date → optional per-entity override for the API LastDate

Output columns

entity_id, date, minimum, maximum, daily_gdd, cumulated_gdd

get_new_token 🔗

get_new_token()

Refresh the authentication token.

setup_gdd_parameters 🔗

setup_gdd_parameters(provider='GLOBAL1', lower_threshold=10.0, upper_threshold=None, start_date=None, end_date=None, reset_cumulative_every_year=False, extrapolate_forecast_data=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for GDD extraction.

Parameters:

Name Type Description Default
provider str

Weather data provider (default 'GLOBAL1').

'GLOBAL1'
lower_threshold float

Lower GDD temperature threshold (required by API).

10.0
upper_threshold float

Upper GDD temperature threshold.

None
start_date str

Start date (YYYY-MM-DD). Can be overwritten per entity by start_date.

None
end_date str

Last date (YYYY-MM-DD). Can be overwritten per entity by end_date.

None
reset_cumulative_every_year bool

Reset the cumulative GDD at the start of each calendar year.

False
extrapolate_forecast_data bool

Extrapolate forecast data when the range extends beyond observations.

False
partial_frequency int

How often to save partial bulk results.

50
column_mapping dict

Custom entity column mapping.

None
output_mapping / exclude_columns / output_columns

Output formatting.

required
use_cache bool

Per-call cache override.

None

get_gdd 🔗

get_gdd(entity_data)

Call the GDD endpoint for one entity.

Resolves StartDate / LastDate from the entity row first, then falls back to the extractor params.

Parameters:

Name Type Description Default
entity_data dict | Series

Must contain id and geometry. May carry start_date / end_date overrides.

required

Returns:

Name Type Description
dict

Raw API response (GrowingDegreeDayResult).

get_gdd_safe 🔗

get_gdd_safe(entity_data) -> dict

Safe wrapper around :meth:get_gdd.

Returns:

Name Type Description
dict dict

{success, data, error, entity_id}.

format_gdd_json 🔗

format_gdd_json(response_json, entity_data=None)

Format the API response into a long DataFrame (one row per day).

Columns

entity_id (if provided), date, minimum, maximum, daily_gdd, cumulated_gdd.

Parameters:

Name Type Description Default
response_json dict

GrowingDegreeDayResult API response.

required
entity_data dict

Original entity row (for id column).

None

Returns:

Type Description

pd.DataFrame: long DataFrame, or empty DataFrame if no data.

process_single_entity_gdd 🔗

process_single_entity_gdd(row, params=None)

Process GDD extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict | Series | DataFrame

Entity data.

required
params dict

Override for self.gdd_params.

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}.

process_entity_gdd_bulk_parallel 🔗

process_entity_gdd_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='gdd', generate_report=False, report_options=None, use_cache=None, spatial_grouping=False, spatial_precision=DEFAULT_SPATIAL_PRECISION, spatial_max_window_days=DEFAULT_MAX_WINDOW_DAYS)

Bulk GDD extraction with threading, partial saves, and merge logic.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities with at least id and geometry; may provide per-row start_date / end_date overrides.

required
spatial_grouping bool

If True, group fields by geohash(centroid) × provider, call the API once per group over the union of the group's date windows, and slice each member out of that single response. GDD depends only on the centroid + date window + thresholds, so neighbouring fields dedup safely. cumulated_gdd accumulates from the request start, so each member's series is re-zeroed to its own start date (year by year when reset_cumulative_every_year is set) — the output matches an ungrouped run row for row. With use_cache, the geohash-keyed spatial cache is consulted per cell before any request.

False
spatial_precision int

Geohash precision for the bucket (default DEFAULT_SPATIAL_PRECISION). Lower = coarser cells = more dedup.

DEFAULT_SPATIAL_PRECISION
spatial_max_window_days int

Cap on a group's widened window (default DEFAULT_MAX_WINDOW_DAYS). Cells whose union exceeds it are split into sub-groups, so one long-history field cannot force a decade-long pull on every field sharing its cell.

DEFAULT_MAX_WINDOW_DAYS

GDDOffsetExtractor 🔗

Bases: BaseExtractor

Extracts GDD-offset (Growing Degree Days threshold offset) analytics.

Given a base date, a lower/upper temperature threshold and a list of GDD offsets, the API returns the forecast date at which each cumulative GDD offset is reached for the point location (field centroid). The output is a wide DataFrame: one row per entity with one column per requested offset (offset_<N> → reached date).

Endpoint: {weather_url}/analytics/gdd-offset

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_GDDOffset.ipynb

Args (setup_gdd_offset_parameters): provider (str): Weather data provider. Default: 'GLOBAL1' lower_threshold (float): Lower GDD temperature threshold. Default: 10.0 upper_threshold (float): Upper GDD temperature threshold. Default: 30.0 date (str): Base date in YYYY-MM-DD the offsets are measured from (maps to API Date). Default: None offsets (list[int]): Cumulative GDD offsets to resolve to reached dates. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required) start_date → optional per-entity override for the API Date offsets → optional per-entity override for the GDD offsets list (int, comma-separated string, or list)

Output columns

entity_id, offset_ (one column per requested offset; value = date the offset is reached, YYYY-MM-DD)

get_new_token 🔗

get_new_token()

Refresh the authentication token.

setup_gdd_offset_parameters 🔗

setup_gdd_offset_parameters(provider='GLOBAL1', lower_threshold=10.0, upper_threshold=30.0, date=None, offsets=None, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for GDD offset extraction.

Parameters:

Name Type Description Default
provider str

Weather data provider (default 'GLOBAL1').

'GLOBAL1'
lower_threshold float

Lower GDD temperature threshold.

10.0
upper_threshold float

Upper GDD temperature threshold.

30.0
date str

Base date (YYYY-MM-DD). Default None — can be overwritten per entity by start_date.

None
offsets int | list[int] | str

GDD offsets to query. Default None — can be overwritten per entity by the offsets column (int, list, or comma-separated string).

None
partial_frequency int

How often to save partial bulk results.

50
column_mapping dict

Custom entity column mapping.

None
output_mapping / exclude_columns / output_columns

Output formatting.

required
use_cache bool

Per-call cache override.

None

get_gdd_offset 🔗

get_gdd_offset(entity_data)

Call the GDD-offset endpoint for one entity.

Resolves Date and Offsets from the entity row first, then falls back to the extractor params.

Parameters:

Name Type Description Default
entity_data dict | Series

Must contain id and geometry. May carry start_date / date and offsets overrides.

required

Returns:

Type Description

list[dict]: Raw API response (one element per offset).

get_gdd_offset_safe 🔗

get_gdd_offset_safe(entity_data) -> dict

Safe wrapper around :meth:get_gdd_offset.

Returns:

Name Type Description
dict dict

{success, data, error, entity_id}.

format_gdd_offset_json 🔗

format_gdd_offset_json(response_json, entity_data=None)

Format the API response into a single-row wide DataFrame.

Columns

entity_id (if provided), offset_<N> for each offset returned (value = date at which offset is reached, YYYY-MM-DD).

Parameters:

Name Type Description Default
response_json list[dict]

API response.

required
entity_data dict

Original entity row (for id column).

None

Returns:

Type Description

pd.DataFrame: single-row wide frame, or empty DataFrame if no data.

process_single_entity_gdd_offset 🔗

process_single_entity_gdd_offset(row, params=None)

Process GDD offset extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict | Series | DataFrame

Entity data.

required
params dict

Override for self.gdd_offset_params.

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}.

process_entity_gdd_offset_bulk_parallel 🔗

process_entity_gdd_offset_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='gdd_offset', generate_report=False, report_options=None, use_cache=None, spatial_grouping=False, spatial_precision=DEFAULT_SPATIAL_PRECISION)

Bulk GDD offset extraction with threading, partial saves, and merge logic.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities with at least id and geometry; may provide per-row start_date/date and offsets overrides.

required
spatial_grouping bool

If True, group fields by geohash(centroid) × start_date × offsets × provider, call the API once per group and broadcast. The forecast reached-date is a function of the centroid, base date, thresholds and offsets — safe to dedup across neighbours.

False
spatial_precision int

Geohash precision for the bucket (default DEFAULT_SPATIAL_PRECISION).

DEFAULT_SPATIAL_PRECISION