Crop development and stressors extractors🔗
Generated from the source docstrings by mkdocstrings at build time.
See the API reference overview for the other groups.
DiseaseExtractor 🔗
Bases: BaseExtractor
Extracts crop disease risk analytics for agricultural entities.
Evaluates disease risk based on weather conditions and crop type. Returns daily disease parameter values (e.g., infection risk, sporulation, severity) for supported crops (corn, soybeans).
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_disease.ipynb
Args (setup_disease_parameters): start_date (str): Start date in YYYY-MM-DD format. Default: None end_date (str): End date in YYYY-MM-DD format. Default: None kpi_filter (dict): KPI aggregation rules applied to the daily disease series. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); crop, start_date, end_date (optional overrides)
Output columns
entity_id, date, + dynamic disease parameter columns (flattened from API response)
get_new_token 🔗
Implements token refresh logic for DiseaseExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_disease_parameters 🔗
setup_disease_parameters(start_date=None, end_date=None, partial_frequency=50, kpi_filter=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure params for disease extraction for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
start_date
|
str
|
Default start date (YYYY-MM-DD) used when entity row has no start_date |
None
|
end_date
|
str
|
Default end date (YYYY-MM-DD) used when entity row has no end_date |
None
|
partial_frequency
|
int
|
How often to save partial results |
50
|
kpi_filter
|
dict
|
KPI aggregation config. See api_utils.validate_kpi_filter for the supported aggregations and their threshold/window contracts. |
None
|
column_mapping
|
dict
|
Override input column mapping for this extractor. Example: {"id": "entity_id", "geometry": "wkt", "crop": "crop_type"} |
None
|
get_disease_data 🔗
Get disease data for selected parameters using POST request
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data with id, geometry, start_date, end_date |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Disease API response |
get_disease_data_safe 🔗
Safe wrapper around get_disease_data(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain 'id', 'geometry', 'start_date', 'end_date' keys |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str |
dict
|
} |
format_disease_json 🔗
Format disease API response into a DataFrame. Dynamically handles all disease parameters returned by the API.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_disease_json
|
list or dict
|
JSON response from get_disease_data |
required |
entity_data
|
dict
|
Original entity data to include ID in output |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: DataFrame with date and all disease parameter columns |
process_single_entity_disease 🔗
Process a single entity for disease data extraction with optional KPI filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
pd.DataFrame, pd.Series, or dict
|
Entity data with required fields |
required |
params
|
dict
|
Override disease_params |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains 'data' (DataFrame) and 'error' |
process_entity_disease_bulk_parallel 🔗
process_entity_disease_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='disease', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity disease data using threads + progress bar, with optional fail-safe retry and filter capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process |
required |
params
|
dict
|
Disease parameters override |
None
|
max_workers
|
int
|
Number of threads to use |
5
|
output_path
|
str
|
Directory to save results |
None
|
partial_frequency
|
int
|
How often to save partial results |
50
|
fail_safe
|
bool
|
If True, retries only previously failed entities |
False
|
filter_column
|
str
|
Column name to filter entities by |
None
|
filter_value
|
any
|
Value to filter |
None
|
filter_type
|
str
|
'exclude' or 'include' |
'exclude'
|
merge_existing
|
str
|
Merge strategy |
None
|
skip_export
|
bool
|
If True, skip final export |
False
|
prefix
|
str
|
Prefix for output filenames |
'disease'
|
use_cache
|
bool
|
Enable/disable caching for this call |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics |
EmergenceExtractor 🔗
Bases: BaseExtractor
Detects crop emergence dates from imagery derived vegetation index time series.
For each field (id + geometry + crop) returns an estimated emergence date
with a confidence score and status. Supports INSEASON, HISTORICAL and DELAY
detection modes and LR / MR data sources. The season window comes from
setup (season_*, year), never from a per-entity sowing date.
Documentation: https://docs.earthdaily.com/agro/library/Emergence/ Notebook: EDAgriculture_Emergence_Processor_Function_Dev.ipynb API endpoint: {agro_urls["emergence_url"][env]}/... (env-dependent)
Args (setup_emergence_parameters): emergence_type (str): Detection type ('INSEASON', 'HISTORICAL', 'DELAY'). Default: 'INSEASON' season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 data_source (str): Data source ('LR', 'MR'). Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
For HISTORICAL, a per-field historical_seasons input (via column_mapping) lists the
calendar years a field actually grew the crop, as a comma-separated string
("2020,2022,2024") or a list of ints. When supplied, emergence_year_N is kept only for
those years (others nulled) and a separate avg_emergence_matching_years column holds
the average recomputed over the retained years (MM-DD). historical_average_emergence
(the raw API average over all 5 seasons) is left untouched.
Entity fields (via column_mapping): id, geometry, crop (required); historical_seasons (optional — HISTORICAL only). sowing_date is NOT read: the season window comes from setup.
Output columns
entity_id, emergence_date, confidence, status, emergence_year_1..5, historical_average_emergence, avg_emergence_matching_years (MM-DD)
get_new_token 🔗
Implements token refresh logic for EmergenceExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
validate_crop
staticmethod
🔗
Validate and normalize the crop type for emergence requests.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop
|
str
|
The crop name to validate. |
required |
available_crops
|
list
|
List of accepted crop names (expected in uppercase). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the crop is missing, invalid, or not a string. |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The validated crop name, normalized to uppercase. |
setup_emergence_parameters 🔗
setup_emergence_parameters(emergence_type='INSEASON', season_duration=120, season_start_day=1, season_start_month=4, year=2025, data_source='LR', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for Emergence extraction
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
emergence_type
|
str
|
Selection between INSEASON, HISTORICAL or DELAY |
'INSEASON'
|
season_duration
|
int
|
Crop cycle duration in days |
120
|
season_start_day
|
int
|
Crop cycle start day (1-31) |
1
|
season_start_month
|
int
|
Crop cycle start month (1-12) |
4
|
year
|
int
|
Year to analyze |
2025
|
data_source
|
str
|
Imagery type - 'LR' (low resolution) or 'MR' (medium resolution) |
'LR'
|
publish_af
|
bool
|
Whether to publish to AF (includes id in payload if True). Default: False |
False
|
partial_frequency
|
int
|
How often to save partial results |
50
|
get_emergence_api 🔗
Request Emergence for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
entity data (id, geometry, crop) |
required |
get_emergence_api_safe 🔗
Safe wrapper around get_emergence_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_emergence_json 🔗
Normalize Emergence API response into a clean pandas DataFrame. Supports multiple emergence types: - INSEASON - HISTORICAL - DELAY
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_emergence_json
|
dict
|
API response JSON with keys 'id' and 'data'. |
required |
params
|
dict
|
Emergence parameters (defaults to self.emergence_params). |
None
|
historical_seasons
|
str | list | None
|
HISTORICAL only — the calendar
years a field actually grew the target crop, as a comma-separated string
("2020,2022,2024") or a list of ints. When supplied, emergence_year_N is kept
only for the retained years (others set to None) and a separate
|
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Flat normalized emergence data. Columns depend on emergence_type. |
process_single_entity_emergence 🔗
Process emergence extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data containing id, geometry, and crop |
required |
params
|
dict
|
Emergence parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_emergence_bulk_extraction_parallel 🔗
process_emergence_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='emergence', generate_report=False, report_options=None, use_cache=None)
Bulk processing of emergence requests using threads + progress bar, with optional fail-safe retry, partial save capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id', 'geometry', and 'crop'). |
required |
params
|
dict
|
Override monitoring parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "emergence". |
'emergence'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] } |
GreennessExtractor 🔗
Bases: BaseExtractor
Extracts crop greenness (vegetation vigor) analytics for agricultural entities.
Monitors crop canopy development and vegetation vigor through greenness detection. Evaluates crop health status based on satellite-derived vegetation indices.
Documentation: https://docs.earthdaily.com/agro/library/greenness/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_greenness.ipynb
Args (setup_greenness_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 sowing_date (str): Sowing date in YYYY-MM-DD format. Default: '2025-04-01' data_source (str): Data source ('LR', 'MR'). Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); crop, sowing_date (optional)
Output columns
entity_id, greenness_date, confidence, status, + detection metadata
get_new_token 🔗
Implements token refresh logic for GreennessExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
validate_crop
staticmethod
🔗
Validate and normalize the crop type for greenness requests.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop
|
str
|
The crop name to validate. |
required |
available_crops
|
set
|
Set of accepted crop names (expected in uppercase). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the crop is missing, invalid, or not a string. |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The validated crop name, normalized to uppercase. |
setup_greenness_parameters 🔗
setup_greenness_parameters(season_duration=120, season_start_day=1, season_start_month=4, year=2025, sowing_date='2025-04-01', data_source='LR', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for greenness detection extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
season_duration
|
int
|
Crop cycle duration in days |
120
|
season_start_day
|
int
|
Crop cycle start day (1-31) |
1
|
season_start_month
|
int
|
Crop cycle start month (1-12) |
4
|
year
|
int
|
Year to analyze |
2025
|
sowing_date
|
str
|
Sowing date in YYYY-MM-DD format |
'2025-04-01'
|
data_source
|
str
|
Imagery type - 'LR' (low resolution) or 'MR' (medium resolution) |
'LR'
|
publish_af
|
bool
|
Whether to publish to AF (includes id in payload if True). Default: False |
False
|
partial_frequency
|
int
|
How often to save partial results |
50
|
get_greenness_api 🔗
Request greenness detection for an entity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
entity data (id, geometry, crop) |
required |
get_greenness_api_safe 🔗
Safe wrapper around get_greenness_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_greenness_json 🔗
Normalize Greenness Detection API response into a clean pandas DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_greenness_json
|
dict
|
API response JSON. |
required |
params
|
dict
|
Greenness parameters (defaults to self.greenness_params). |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Flat normalized greenness data. |
process_single_entity_greenness 🔗
Process greenness extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data containing id, geometry, and crop |
required |
params
|
dict
|
Greenness parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_greenness_bulk_extraction_parallel 🔗
process_greenness_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='greenness', generate_report=False, report_options=None, use_cache=None)
Bulk processing of greenness requests using threads + progress bar, with optional fail-safe retry, partial save capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id', 'geometry', and 'crop'). |
required |
params
|
dict
|
Override greenness parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "greenness". |
'greenness'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] } |
HarvestExtractor 🔗
Bases: BaseExtractor
Extracts harvest detection analytics for agricultural entities.
Identifies harvest dates and monitors harvest progress using satellite-based temporal analysis. Supports in-season, historical and harvest-readiness detection.
Documentation: https://docs.earthdaily.com/agro/library/Harvest_Detection/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_harvest.ipynb
Args (setup_harvest_parameters): harvest_type (str): Detection type - 'INSEASON_HARVEST', 'HISTORICAL_HARVEST' or 'HARVEST_READINESS'. Default: 'INSEASON_HARVEST' season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 data_source (str): Data source ('LR', 'MR'). Default: 'LR' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
For HISTORICAL_HARVEST, a per-field historical_seasons input (via column_mapping) lists
the calendar years a field actually grew the crop, as a comma-separated string
("2020,2022,2024") or a list of ints. When supplied, harvest_year_N is kept only for those
years (others nulled) and a separate avg_harvest_matching_years column holds the average
recomputed over the retained years (MM-DD). historical_harvest_average (the raw API
average over all 5 seasons) is left untouched.
Entity fields (via column_mapping): id, geometry, crop (required); historical_seasons (optional — HISTORICAL_HARVEST only). sowing_date is NOT read: the season window comes from setup.
Output columns
Varies by harvest_type. INSEASON_HARVEST: entity_id, harvest_date, harvest_status HISTORICAL_HARVEST: entity_id, harvest_year_1..5, historical_harvest_average, avg_harvest_matching_years (MM-DD) HARVEST_READINESS: entity_id, harvest_readiness_date, is_ready
get_new_token 🔗
Implements token refresh logic for HarvestExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
validate_crop
staticmethod
🔗
Validate and normalize the crop type for harvest requests.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop
|
str
|
The crop name to validate. |
required |
available_crops
|
list
|
List of accepted crop names (expected in uppercase). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the crop is missing, invalid, or not a string. |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The validated crop name, normalized to uppercase. |
setup_harvest_parameters 🔗
setup_harvest_parameters(harvest_type='INSEASON_HARVEST', season_duration=120, season_start_day=1, season_start_month=4, year=2025, data_source='LR', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for harvest extraction
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
harvest_type
|
str
|
One of 'INSEASON_HARVEST' (harvest date + status for the current season), 'HISTORICAL_HARVEST' (harvest date for each of the last 5 seasons) or 'HARVEST_READINESS' (readiness date + is_ready flag) |
'INSEASON_HARVEST'
|
season_duration
|
int
|
Crop cycle duration in days |
120
|
season_start_day
|
int
|
Crop cycle start day (1-31) |
1
|
season_start_month
|
int
|
Crop cycle start month (1-12) |
4
|
year
|
int
|
Year to analyze |
2025
|
data_source
|
str
|
Imagery type - 'LR' (low resolution) or 'MR' (medium resolution) |
'LR'
|
publish_af
|
bool
|
Whether to publish to AF (includes id in payload if True). Default: False |
False
|
partial_frequency
|
int
|
How often to save partial results |
50
|
get_harvest_api 🔗
Request harvest for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
entity data (id, geometry, crop) |
required |
get_harvest_api_safe 🔗
Safe wrapper around get_harvest_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_harvest_json 🔗
Normalize Harvest API response into a clean pandas DataFrame. Supports multiple harvest types: - INSEASON_HARVEST - HISTORICAL_HARVEST - HARVEST_READINESS
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_harvest_json
|
dict
|
API response JSON with keys 'id' and 'data'. |
required |
params
|
dict
|
Harvest parameters (defaults to self.harvest_params). |
None
|
historical_seasons
|
str | list | None
|
HISTORICAL_HARVEST only — the
calendar years a field actually grew the target crop, as a comma-separated
string ("2020,2022,2024") or a list of ints. When supplied, harvest_year_N is
kept only for the retained years (others set to None) and a separate
|
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Flat normalized harvest data. Columns depend on harvest_type. |
process_single_entity_harvest 🔗
Process harvest extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data containing id, geometry, and crop |
required |
params
|
dict
|
Harvest parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_harvest_bulk_extraction_parallel 🔗
process_harvest_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='harvest', generate_report=False, report_options=None, use_cache=None)
Bulk processing of harvest requests using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id', 'geometry', and 'crop'). |
required |
params
|
dict
|
Override monitoring parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "harvest". |
'harvest'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts. |
PlantedExtractor 🔗
Bases: BaseExtractor
Extracts planted area estimation analytics for agricultural entities.
Estimates planted acreage using crop identification and emergence data. Validates planting status and computes planted area percentages.
Documentation: https://docs.earthdaily.com/agro/library/Planted_Area/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_coverage.ipynb
Args (setup_planted_parameters):
processor_mode (str): Processing mode - 'PLANTED_AREA' (planted area and percentage)
or 'CONTROL' (compares the field against control_threshold). Default: 'PLANTED_AREA'
emergence_date (str): Default emergence date in YYYY-MM-DD. Default: None
(entities may override via the emergence_date row column)
threshold (int): Decision threshold in days. Default: 120
control_threshold (int | float): Percentage threshold, CONTROL mode only. Sent to the
API divided by 100, so the default 4 becomes controlThreshold=0.04. Default: 4
use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); emergence_date (optional per-entity override of params.emergence_date — wire the EmergenceExtractor output in via column_mapping); crop, sowing_date (optional)
Output columns
Varies by processor_mode. PLANTED_AREA: entity_id, planted_area_m2, planted_percentage CONTROL: entity_id, difference, control_threshold, control_result
get_new_token 🔗
Implements token refresh logic for PlantedExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_planted_parameters 🔗
setup_planted_parameters(processor_mode='PLANTED_AREA', emergence_date=None, threshold=120, control_threshold=4, publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for planted area extraction
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
processor_mode
|
str
|
Selection between 'PLANTED_AREA' or 'CONTROL' |
'PLANTED_AREA'
|
emergence_date
|
str
|
Default emergence date in YYYY-MM-DD.
May be overridden per entity via the |
None
|
threshold
|
int
|
Days between images before emergence and second image (default: 120) |
120
|
control_threshold
|
float
|
Percentage threshold for CONTROL mode (default: 4, used as 0.04) |
4
|
publish_af
|
bool
|
Whether to publish to AF (includes id in payload if True). Default: False |
False
|
partial_frequency
|
int
|
How often to save partial results |
50
|
get_planted_api 🔗
Request planted area calculation for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
entity data (id, geometry) |
required |
get_planted_api_safe 🔗
Safe wrapper around get_planted_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str |
dict
|
} |
format_planted_json 🔗
Normalize Planted Area API response into a clean pandas DataFrame. Supports two processor modes: - PLANTED_AREA: Returns planted area in m² and percentage - CONTROL: Returns difference, control threshold, and result boolean
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_planted_json
|
dict
|
API response JSON. |
required |
params
|
dict
|
Planted area parameters (defaults to self.planted_params). |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Flat normalized planted area data. Columns depend on processor_mode. |
process_single_entity_planted 🔗
Process planted area extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data containing id and geometry |
required |
params
|
dict
|
Planted area parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_planted_bulk_extraction_parallel 🔗
process_planted_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='planted', generate_report=False, report_options=None, use_cache=None)
Bulk processing of planted area requests using threads + progress bar, with optional fail-safe retry and partial save capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id' and 'geometry'). |
required |
params
|
dict
|
Override planted area parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "planted". |
'planted'
|
use_cache
|
bool
|
Enable/disable caching for this call. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] } |
GDDExtractor 🔗
Bases: BaseExtractor
Extracts GDD (Growing Degree Days) analytics for agricultural entities.
Computes daily and cumulated Growing Degree Days from a start date to a last date, using a lower (and optional upper) temperature threshold. The output is a long DataFrame: one row per entity per day, with columns for minimum/maximum temperature, daily GDD and cumulated GDD.
Endpoint: {weather_url}/analytics/gdd
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgro_growing_degree_days.ipynb
Args (setup_gdd_parameters): provider (str): Weather data provider. Default: 'GLOBAL1' lower_threshold (float): Lower GDD temperature threshold (required by API). Default: 10.0 upper_threshold (float): Optional upper temperature threshold (capped GDD). Default: None start_date (str): Period start in YYYY-MM-DD (maps to API StartDate). Default: None end_date (str): Period end in YYYY-MM-DD (maps to API LastDate). Default: None reset_cumulative_every_year (bool): Reset the cumulated GDD at each year boundary. Default: False extrapolate_forecast_data (bool): Extend the series with forecast data. Default: False use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping):
id, geometry (required)
start_date → optional per-entity override for the API StartDate
end_date → optional per-entity override for the API LastDate
Output columns
entity_id, date, minimum, maximum, daily_gdd, cumulated_gdd
setup_gdd_parameters 🔗
setup_gdd_parameters(provider='GLOBAL1', lower_threshold=10.0, upper_threshold=None, start_date=None, end_date=None, reset_cumulative_every_year=False, extrapolate_forecast_data=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for GDD extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Weather data provider (default |
'GLOBAL1'
|
lower_threshold
|
float
|
Lower GDD temperature threshold (required by API). |
10.0
|
upper_threshold
|
float
|
Upper GDD temperature threshold. |
None
|
start_date
|
str
|
Start date ( |
None
|
end_date
|
str
|
Last date ( |
None
|
reset_cumulative_every_year
|
bool
|
Reset the cumulative GDD at the start of each calendar year. |
False
|
extrapolate_forecast_data
|
bool
|
Extrapolate forecast data when the range extends beyond observations. |
False
|
partial_frequency
|
int
|
How often to save partial bulk results. |
50
|
column_mapping
|
dict
|
Custom entity column mapping. |
None
|
output_mapping / exclude_columns / output_columns
|
Output formatting. |
required | |
use_cache
|
bool
|
Per-call cache override. |
None
|
get_gdd 🔗
Call the GDD endpoint for one entity.
Resolves StartDate / LastDate from the entity row first, then
falls back to the extractor params.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict | Series
|
Must contain |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Raw API response ( |
get_gdd_safe 🔗
Safe wrapper around :meth:get_gdd.
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
|
format_gdd_json 🔗
Format the API response into a long DataFrame (one row per day).
Columns
entity_id (if provided), date, minimum, maximum,
daily_gdd, cumulated_gdd.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
|
required |
entity_data
|
dict
|
Original entity row (for id column). |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: long DataFrame, or empty DataFrame if no data. |
process_single_entity_gdd 🔗
Process GDD extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict | Series | DataFrame
|
Entity data. |
required |
params
|
dict
|
Override for |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
|
process_entity_gdd_bulk_parallel 🔗
process_entity_gdd_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='gdd', generate_report=False, report_options=None, use_cache=None, spatial_grouping=False, spatial_precision=DEFAULT_SPATIAL_PRECISION, spatial_max_window_days=DEFAULT_MAX_WINDOW_DAYS)
Bulk GDD extraction with threading, partial saves, and merge logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities with at least |
required |
spatial_grouping
|
bool
|
If True, group fields by |
False
|
spatial_precision
|
int
|
Geohash precision for the bucket (default
|
DEFAULT_SPATIAL_PRECISION
|
spatial_max_window_days
|
int
|
Cap on a group's widened window (default
|
DEFAULT_MAX_WINDOW_DAYS
|
GDDOffsetExtractor 🔗
Bases: BaseExtractor
Extracts GDD-offset (Growing Degree Days threshold offset) analytics.
Given a base date, a lower/upper temperature threshold and a list of GDD
offsets, the API returns the forecast date at which each cumulative GDD
offset is reached for the point location (field centroid). The output is
a wide DataFrame: one row per entity with one column per requested
offset (offset_<N> → reached date).
Endpoint: {weather_url}/analytics/gdd-offset
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_GDDOffset.ipynb
Args (setup_gdd_offset_parameters): provider (str): Weather data provider. Default: 'GLOBAL1' lower_threshold (float): Lower GDD temperature threshold. Default: 10.0 upper_threshold (float): Upper GDD temperature threshold. Default: 30.0 date (str): Base date in YYYY-MM-DD the offsets are measured from (maps to API Date). Default: None offsets (list[int]): Cumulative GDD offsets to resolve to reached dates. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping):
id, geometry (required)
start_date → optional per-entity override for the API Date
offsets → optional per-entity override for the GDD offsets list
(int, comma-separated string, or list)
Output columns
entity_id, offset_
setup_gdd_offset_parameters 🔗
setup_gdd_offset_parameters(provider='GLOBAL1', lower_threshold=10.0, upper_threshold=30.0, date=None, offsets=None, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for GDD offset extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Weather data provider (default |
'GLOBAL1'
|
lower_threshold
|
float
|
Lower GDD temperature threshold. |
10.0
|
upper_threshold
|
float
|
Upper GDD temperature threshold. |
30.0
|
date
|
str
|
Base date ( |
None
|
offsets
|
int | list[int] | str
|
GDD offsets to query.
Default |
None
|
partial_frequency
|
int
|
How often to save partial bulk results. |
50
|
column_mapping
|
dict
|
Custom entity column mapping. |
None
|
output_mapping / exclude_columns / output_columns
|
Output formatting. |
required | |
use_cache
|
bool
|
Per-call cache override. |
None
|
get_gdd_offset 🔗
Call the GDD-offset endpoint for one entity.
Resolves Date and Offsets from the entity row first, then
falls back to the extractor params.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict | Series
|
Must contain |
required |
Returns:
| Type | Description |
|---|---|
|
list[dict]: Raw API response (one element per offset). |
get_gdd_offset_safe 🔗
Safe wrapper around :meth:get_gdd_offset.
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
|
format_gdd_offset_json 🔗
Format the API response into a single-row wide DataFrame.
Columns
entity_id (if provided), offset_<N> for each offset
returned (value = date at which offset is reached, YYYY-MM-DD).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
list[dict]
|
API response. |
required |
entity_data
|
dict
|
Original entity row (for id column). |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: single-row wide frame, or empty DataFrame if no data. |
process_single_entity_gdd_offset 🔗
Process GDD offset extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict | Series | DataFrame
|
Entity data. |
required |
params
|
dict
|
Override for |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
|
process_entity_gdd_offset_bulk_parallel 🔗
process_entity_gdd_offset_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='gdd_offset', generate_report=False, report_options=None, use_cache=None, spatial_grouping=False, spatial_precision=DEFAULT_SPATIAL_PRECISION)
Bulk GDD offset extraction with threading, partial saves, and merge logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities with at least |
required |
spatial_grouping
|
bool
|
If True, group fields by |
False
|
spatial_precision
|
int
|
Geohash precision for the bucket (default
|
DEFAULT_SPATIAL_PRECISION
|