Foundational extractors🔗
Generated from the source docstrings by mkdocstrings at build time.
See the API reference overview for the other groups.
CoverageExtractor 🔗
Bases: BaseExtractor
Extracts satellite image coverage analytics for agricultural entities.
Retrieves available satellite imagery metadata (coverage percentage, sensor, date, mask) for a given geometry and date range. Supports duplicate filtering across sensors and multi-sensor coverage analysis.
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_coverage.ipynb
Args (setup_coverage_parameters): vegetation_index (str): Vegetation index type ('NDVI', 'EVI', etc.). Default: 'NDVI' start_date (str): Start date in YYYY-MM-DD format. Default: '2025-01-01' end_date (str): End date in YYYY-MM-DD format. Default: None — no upper bound, i.e. every image from start_date onward is returned (the query becomes image.date=$gte:start_date). It is not defaulted to today; set it explicitly to bound the catalogue query. clear_cover_min (int): Minimum clear sky coverage percentage (0-100). Default: 90 clear_cover_max (int): Maximum clear sky coverage percentage (0-100). Default: 100 use_specific_date (bool): Restrict to exact date. Default: False filter (str): Filtering mode - 'none', 'duplicate' or 'crop_coverage'. Default: 'none' 'duplicate' collapses near-simultaneous multi-sensor captures (see delay). 'crop_coverage' expands the query per entity across prior seasons and requires historical_seasons; the per-entity window is computed from the entity's start_date/end_date month-day applied to each historical year. delay (int): Max days between multi-sensor captures for duplicate filter. Default: 3 mask (str): Cloud mask type — 'Auto', 'Native', 'ACM', 'ML', 'MLCirrus' or 'All' (case-insensitive). 'All' returns one row per image per mask, so the same image appears several times and the 'mask' output column tells them apart — use it to compare masks. Not compatible with filter='duplicate'. Default: 'auto' recalibration (bool): Apply cross-sensor recalibration to coverage metrics. Default: False historical_seasons (list[int]): Prior season years used when filter='crop_coverage' (required for that filter, e.g. [2024, 2023, 2022]). Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required)
Output columns
image_id, coverage_percent, date, mask, sensor, spatial_resolution
get_new_token 🔗
Implements token refresh logic for CoverageExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_coverage_parameters 🔗
setup_coverage_parameters(vegetation_index='NDVI', start_date='2025-01-01', end_date=None, clear_cover_min=90, clear_cover_max=100, use_specific_date=False, filter='none', delay=3, mask='auto', partial_frequency=50, recalibration=False, historical_seasons=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure params for imagery catalog extraction for an entity.
When filter='crop_coverage', historical_seasons must be provided. The extraction period is computed per entity from start_date/end_date fields across all historical years (smallest year start to largest year end).
get_satellite_coverage_by_geometry 🔗
Query the coverage API to get imagery metadata for a given entity geometry. If recalibration is enabled in coverage_params, uses the recalibration endpoint and includes crop in the payload.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. If recalibration=True, must also contain a 'crop' key. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
API response with imagery metadata |
get_satellite_coverage_by_geometry_safe 🔗
Safe wrapper around get_satellite_coverage_by_geometry(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_coverage_json 🔗
Process API response to extract image ID, coverage percent, formatted date, sensor, spatial resolution, and mask information.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_coverage_json
|
str | list | dict
|
API response (stringified JSON, parsed list, or dict). |
required |
entity_id
|
str
|
Entity ID for logging purposes. |
'unknown'
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Cleaned DataFrame of imagery coverage. |
process_single_entity_coverage 🔗
Process coverage extraction for a single entity with retry logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data containing id and geometry |
required |
params
|
dict
|
Coverage parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_entity_coverage_bulk_parallel 🔗
process_entity_coverage_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='coverage', flm_params=None, recalibration=False, generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity coverage using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain 'id' and 'geometry'. |
required |
params
|
dict
|
Coverage parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "coverage". |
'coverage'
|
flm_params
|
dict
|
FLM parameters. If provided, sets recalibration=True and stores on self. |
None
|
recalibration
|
bool
|
If True, uses recalibration API endpoint. Default: False. |
False
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts. |
FLMExtractor 🔗
Bases: BaseExtractor
Extracts Field Level Maps (vegetation index maps) for agricultural entities.
Generates field-level vegetation index maps (NDVI, EVI, NDRE, etc.) from satellite imagery. Supports map statistics extraction, direct download links, and raster file export with configurable clipping and buffering.
Documentation: https://docs.earthdaily.com/agro/library/Field%20Level%20Maps/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_FLM_extraction_functions.ipynb
Args (setup_flm_parameters): vegetation_index (str): Index type ('NDVI', 'EVI', 'NDRE', etc.). Default: 'NDVI' map_format (str): Output format (None, 'png', 'tiff'). Default: None output_epsg (int): Output coordinate system. Default: 4326 postprocess (str): Post-processing mode ('stats', 'links', 'file', 'histogram'). Default: 'stats' 'histogram' sets histogram=true on the query and returns the per-bucket breakdown as well as the global min/max/mean. extract_stats (bool): Extract map statistics. Default: False directLinks (bool): Return direct download links. Default: False clipping (str): Clipping mode ('FieldBorder', 'None'). Default: 'FieldBorder' buffer (int): Buffer around field border in meters. Default: 0 number_bins (int): Number of legend bins for the rendered map. Default: None legendType (str): Legend style for the rendered map. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required)
Output columns
Varies by postprocess mode: stats: stat_max, stat_mean, stat_min links: image_png_link, worldfile_link, thumbnail_link, bbox_*, map_width, map_height file: status, map_format, file_count, total_size_bytes, saved_files histogram: stat_min, stat_max, stat_mean, value_min_{i}, value_max_{i}, num_pixels_{i}, area_{i} (one set of {i} columns per bucket, single row per entity)
get_new_token 🔗
Implements token refresh logic for FLMExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_flm_parameters 🔗
setup_flm_parameters(vegetation_index='NDVI', map_format=None, output_epsg=4326, postprocess='stats', skip_existing=True, output_path=None, extract_stats=False, partial_frequency=50, directLinks=False, clipping='FieldBorder', buffer=0, number_bins=None, legendType=None, use_cache=None)
Configure parameters for FLM extraction
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
vegetation_index
|
str
|
Vegetation index to extract |
'NDVI'
|
map_format
|
str
|
Format of the map ('png', 'tiff.zip', 'shp.zip') |
None
|
output_epsg
|
int
|
Output EPSG code for projection |
4326
|
postprocess
|
str
|
Post-processing type - 'stats', 'file', 'links', or 'histogram'. Default: 'stats' |
'stats'
|
skip_existing
|
bool
|
Skip download if file already exists |
True
|
output_path
|
str
|
Path to save downloaded maps |
None
|
extract_stats
|
bool
|
If True, extract statistics only (no file download) |
False
|
partial_frequency
|
int
|
How often to save partial results |
50
|
directLinks
|
bool
|
Whether to use direct links in API response. Default: False Note: Automatically set to True when postprocess='links' |
False
|
clipping
|
str
|
Clipping method - 'FieldBorder' or 'Bbox'. Default: 'FieldBorder' Note: Only applicable for COLORCOMPOSITION vegetation index |
'FieldBorder'
|
buffer
|
int
|
Buffer distance in meters around geometry. Default: 0 |
0
|
number_bins
|
int
|
Number of histogram bins (1-255). When set,
appends |
None
|
legendType
|
str
|
Legend type - one of 'Fixed', 'Dynamic', 'Common'.
When set, appends |
None
|
get_flm_map 🔗
Process FLM map: either download file OR extract statistics OR get direct links Uses parameters from self.flm_params (set by setup_parameters)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id' and 'geometry' |
required |
image_id
|
str
|
Image ID to download |
required |
file_extension
|
str
|
File extension override |
None
|
Returns:
| Type | Description |
|---|---|
|
requests.Response: Raw API response |
get_flm_map_safe 🔗
Safe wrapper around get_flm_map(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
image_id
|
str
|
Image ID |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str |
dict
|
} |
format_flm_map_stats_json 🔗
Extract legend statistics and metadata from Map Reference JSON. Converts raw response to JSON at the beginning.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_flm_map
|
dict or Response
|
API response JSON with legend, mapSize, bBox, etc. |
required |
range
|
bool
|
If True, also return ranges DataFrame. Default is False. |
False
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame or tuple: - If range=False: stats_df only - If range=True: (stats_df, ranges_df) |
format_flm_map_histogram_json 🔗
Extract histogram stats and per-range values from Map Reference JSON.
Produces a single-row DataFrame per entity:
- global stats: stat_min, stat_max, stat_mean
- per-bucket columns (1-indexed): value_min_{i},
value_max_{i}, num_pixels_{i}, area_{i}
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_flm_map
|
dict or Response
|
API response carrying a
|
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: single-row wide DataFrame. |
format_flm_map_file 🔗
Save FLM map file response (PNG, TIFF.ZIP, or SHP.ZIP) to disk. Handles extraction of zipped files and proper naming.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id' and optional 'name' |
required |
image_id
|
str
|
Image ID for filename |
required |
response
|
Response
|
Raw response from get_flm_map |
required |
output_path
|
str
|
Directory to save files. Uses self.flm_params['output_path'] if not provided. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Status and file path information |
format_flm_links_json 🔗
Extract direct links and metadata from FLM API response. Converts raw response to JSON at the beginning.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_flm_map
|
dict or Response
|
API response JSON with _links, worldFile, etc. |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: DataFrame containing links and metadata (one row per entity) |
process_single_entity_flm 🔗
Process FLM extraction for a single entity with retry logic. Supports four postprocess modes: 'stats', 'links', 'file', or 'histogram'
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data containing id, geometry, and image_id |
required |
params
|
dict
|
FLM parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_entity_flm_bulk_parallel 🔗
process_entity_flm_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='flm', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity FLM maps using threads + progress bar, with optional fail-safe retry and filter capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain 'id', 'geometry', and 'image_id'. |
required |
params
|
dict
|
FLM parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "flm". |
'flm'
|
use_cache
|
bool
|
Enable/disable caching for this call. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. |
VegationTsExtractor 🔗
Bases: BaseExtractor
Extracts high-resolution vegetation time series for agricultural entities.
Retrieves historical and near-real-time vegetation index time series (NDVI, EVI, LAI, etc.) at field level. Supports period-based and target-date extraction modes, KPI aggregation, and multi-year historical comparison.
Documentation: https://docs.earthdaily.com/agro/library/Vegetation_time_series/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_VTS.ipynb
Args (setup_vegetation_ts_parameters): start_date (str): Start date in YYYY-MM-DD format end_date (str): End date in YYYY-MM-DD format vegetation_index (str): Index type ('NDVI', 'EVI', 'CVI', 'GNDVI', 'NDWI', 'LAI'). Default: 'NDVI' is_extrapolated (bool): Include extrapolated values. Default: True limit (int): Max number of records. Default: 3000 historical_years (int): Number of historical years. Default: 10 extraction_mode (str): 'period' or 'target_dates'. Default: 'period' target_dates (list): Specific dates to extract (for target_dates mode) kpi_filter (dict): KPI aggregation config (kpi_name, aggregation, threshold) use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); crop (required for LAI); start_date, end_date (optional overrides)
Output columns
entity_id, date, value (vegetation index), + historical year columns when applicable
get_new_token 🔗
Implements token refresh logic for VegationTsExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_vegetation_ts_parameters 🔗
setup_vegetation_ts_parameters(start_date: str = None, end_date: str = None, vegetation_index='NDVI', is_extrapolated=True, limit=3000, historical_years=10, partial_frequency=50, extraction_mode='period', target_dates=None, kpi_filter=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for vegetation index extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
start_date
|
str
|
Baseline start date in YYYY-MM-DD format. Used as default when entity_data does not provide its own start_date. |
None
|
end_date
|
str
|
Baseline end date in YYYY-MM-DD format. Used as default when entity_data does not provide its own end_date. |
None
|
vegetation_index
|
str
|
Vegetation index to extract (NDVI, EVI) |
'NDVI'
|
is_extrapolated
|
bool
|
Whether to include extrapolated values |
True
|
limit
|
int
|
Maximum number of records to retrieve |
3000
|
historical_years
|
int
|
Number of historical years to include (max 15) |
10
|
partial_frequency
|
int
|
How often to save partial results |
50
|
extraction_mode
|
str
|
Extraction mode. - 'period': seasonal MM-DD window extraction. Uses start_date/end_date as baseline (entity_data can override). Expands by historical_years for KPI comparison. - 'windows': absolute date range extraction. Uses start_date/end_date as baseline (entity_data can override). No historical expansion. - 'specific_dates': punctual extraction at exact target dates. |
'period'
|
target_dates
|
list
|
|
None
|
kpi_filter
|
dict
|
KPI configuration |
None
|
process_single_entity_specific_dates 🔗
Process a single entity for vegetation time series extraction at specific dates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
Series or dict
|
Entity data with 'id' and 'name' |
required |
params
|
dict
|
Override vegetation_ts_params |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains 'data' (DataFrame with columns: date, field_name, ndvi_value) and 'error' |
get_vegetation_api 🔗
Get vegetation time series data from VTS API.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data containing: - 'id' (str): Season field ID - 'start_date' (str): Start date in YYYY-MM-DD format (required for 'period' mode) - 'end_date' (str): End date in YYYY-MM-DD format (required for 'period' mode) Note: For 'windows' mode, start_date and end_date are optional and will use target_dates from vegetation_ts_params if not provided |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
JSON response with vegetation time series data |
get_vegetation_api_safe 🔗
Safe wrapper around get_vegetation_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain 'id', 'start_date', and 'end_date' keys |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str |
dict
|
} |
format_vegetation_ts_json 🔗
Process vegetation API response into a DataFrame with date and value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_vegetation_ts_json
|
list
|
JSON response from get_vegetation_api (list of date/value records) |
required |
entity_data
|
dict
|
Original entity data to include ID in output |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: DataFrame with columns: entity_id (optional), date, value |
process_single_entity_vegetation_ts 🔗
Process a single entity for vegetation time series extraction with optional KPI filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
Series or dict
|
Entity data with required fields: - 'id': Entity identifier - 'geometry': WKT geometry string - 'start_date': Start date (YYYY-MM-DD) - 'end_date': End date (YYYY-MM-DD) - 'years' (optional): List of historical years for KPI comparison |
required |
params
|
dict
|
Override vegetation_ts_params (includes kpi_filter config) |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains 'data' (DataFrame) and 'error' - If KPI requested: DataFrame with metadata + KPI values (one row per entity) - If NO KPI: DataFrame with metadata + daily date/value data (multiple rows) |
process_entity_vegetation_ts_bulk_parallel 🔗
process_entity_vegetation_ts_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='vegetation_ts', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity vegetation time series using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain 'id' and 'geometry'. |
required |
params
|
dict
|
Vegetation time series parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "vegetation_ts". |
'vegetation_ts'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts. |
process_specific_dates_bulk_parallel 🔗
process_specific_dates_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', skip_export=False, prefix='vegetation_specific_dates', generate_report=False, report_options=None, use_cache=None)
Bulk processing of vegetation time series at specific dates using parallel threads, with optional caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities with 'id' and 'name' columns |
required |
params
|
dict
|
Override vegetation_ts_params |
None
|
max_workers
|
int
|
Number of parallel threads |
5
|
output_path
|
str
|
Directory to save results |
None
|
partial_frequency
|
int
|
Save partial results every N entities |
50
|
fail_safe
|
bool
|
Retry previously failed entities |
False
|
filter_column
|
str
|
Column to filter by |
None
|
filter_value
|
any
|
Value to filter |
None
|
filter_type
|
str
|
'exclude' or 'include' |
'exclude'
|
skip_export
|
bool
|
Skip final export |
False
|
prefix
|
str
|
Output filename prefix |
'vegetation_specific_dates'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Results DataFrame, errors, and statistics. When cache is active, also includes cache_hit and cache_miss counts. |
MRTSExtractor 🔗
Bases: BaseExtractor
Extracts Medium Resolution Time Series (MRTS) vegetation data for agricultural entities.
Retrieves multi-sensor vegetation index time series using geometry-based queries. Supports raw and smoothed data extraction, temporal consistency checks, denoising, and end-of-curve extrapolation. Works with Sentinel-2, Landsat-8/9, and other sensors.
Documentation: https://docs.earthdaily.com/agro/library/Vegetation_time_series/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_MRTS_extraction_functions.ipynb
Args (setup_mrts_parameters): start_date (str): Start date in YYYY-MM-DD format. Default: '2025-05-01' end_date (str): End date in YYYY-MM-DD format. Default: '2025-10-15' sensors (list): Sensor list (Sentinel_2, Landsat_8, Landsat_9, etc.). Default: None (all) vegetation_index (str): Index type ('NDVI', 'EVI', 'CVI', 'GNDVI', 'NDWI', 'LAI'). Default: 'NDVI' aggregation (str): Aggregation method ('average', 'median', 'max', 'min'). Default: 'average' smoothing_method (str): Smoothing ('Whittaker', 'SavitzkyGolay', 'None'). Default: 'Whittaker' apply_denoiser (bool): Apply denoising filter. Default: True apply_end_of_curve (bool): Extrapolate end of curve. Default: True clear_cover_min (int): Minimum clear sky percentage (0-100). Default: 100 mask (str): Cloud mask used to discard NotClear pixels ('Native', 'ACM', 'Auto', 'ML', 'MLCirrus'). 'All' is NOT accepted here (coverage-only). Default: 'Auto' — best mask picked per image. Pass None to omit the key entirely (same behaviour). mode (str): Output mode ('full' = raw+smoothed, 'raw' = raw only). Default: 'full' compute_temporal_consistency (bool): Compute consistency checks. Default: True temporal_consistency_threshold (dict): Per-index thresholds. Example: {"Ndvi": 0.06, "Lai": 0.3} output_saturation (bool): Include index saturation flags in the output. Default: True extract_raw_datasets (bool): Include the raw per-image datasets alongside smoothed values. Default: True historical_years (int): Number of prior years to fetch for historical comparison. Default: 10 kpi_filter (dict): KPI aggregation rules applied to the time series. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); crop (required for LAI); start_date, end_date (optional overrides)
Output columns
entity_id, date, raw_value, noised, temporalConsistencyCheck, image_id, smoothed_value (full mode)
__init__ 🔗
Initialize Vegetation Time Series Extractor
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bearer_token
|
str
|
Bearer token for API authentication |
required |
token_expiration
|
datetime
|
Token expiration datetime |
required |
config
|
dict
|
Configuration dictionary with env, paths, etc. |
required |
workflow_ref
|
optional
|
Reference to WorkflowManager for token refresh |
None
|
get_new_token 🔗
Implements token refresh logic for cropidExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_mrts_parameters 🔗
setup_mrts_parameters(start_date='2025-05-01', end_date='2025-10-15', sensors=None, vegetation_index='NDVI', aggregation='average', smoothing_method='Whittaker', apply_denoiser=True, apply_end_of_curve=True, clear_cover_min=100, mask='Auto', output_saturation=True, extract_raw_datasets=True, compute_temporal_consistency=True, temporal_consistency_threshold=None, mode='full', historical_years=10, partial_frequency=50, kpi_filter=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for MRTS (Multi-Resolution Time Series) extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
start_date
|
str
|
Start date of the period (YYYY-MM-DD). Default: '2025-05-01' |
'2025-05-01'
|
end_date
|
str
|
End date of the period (YYYY-MM-DD). Default: '2025-10-15' |
'2025-10-15'
|
sensors
|
list
|
List of sensors to use. Available: Sentinel_2, Landsat_8, Landsat_9, HJ2B_CCD4, GAOFEN_6_WFV3 Default: None (omitted from payload → API uses all available sensors) |
None
|
vegetation_index
|
str
|
Vegetation index ('NDVI', 'EVI', 'CVI', 'GNDVI', 'NDWI', 'LAI'). Default: 'NDVI' |
'NDVI'
|
aggregation
|
str
|
Aggregation method ('average', 'median', 'max', 'min'). Default: 'average' |
'average'
|
smoothing_method
|
str
|
Smoothing method ('Whittaker', 'SavitzkyGolay', 'None'). Default: 'Whittaker' |
'Whittaker'
|
apply_denoiser
|
bool
|
Apply denoising. Default: True |
True
|
apply_end_of_curve
|
bool
|
Apply end of curve extrapolation. Default: True |
True
|
clear_cover_min
|
int
|
Minimum clear sky coverage percentage (0-100). Default: 100 |
100
|
mask
|
str
|
Cloud mask applied before aggregation (TimeSeriesMaskType). One of 'Native', 'ACM', 'Auto', 'ML', 'MLCirrus' (case-insensitive; normalized to the API casing). Unlike the coverage/catalog-imagery endpoint, 'All' is rejected by this endpoint (HTTP 400) — it is a coverage-only value. Default: 'Auto' — the API selects the best available mask per image, which gives the most complete series. This is also what the API does when the key is absent, so the default is explicit, not a behaviour change; pass mask=None to omit it. Notes from live probing (prod, 2026-08-25, 3 sites x 2 seasons): - omitting mask is identical to mask='Auto' (same observations, not just the same count) — it is NOT Native and NOT ML - 'Auto' picks the best mask per image, so a single response can mix Native/ACM/ML labels; the 'mask' field on each raw point reports the mask actually applied, not the one requested - on fields where Auto happens to resolve to one mask throughout, omitted / 'Auto' / that mask are indistinguishable — compare observations, not counts - 'ACM' returns noticeably fewer raw observations than Native/ML - 'MLCirrus' is accepted but returned no data on any field tested |
'Auto'
|
output_saturation
|
bool
|
Include saturation information. Default: True |
True
|
extract_raw_datasets
|
bool
|
Extract raw datasets in addition to smoothed data. Default: True |
True
|
compute_temporal_consistency
|
bool
|
Compute temporal consistency metrics. Default: True |
True
|
temporal_consistency_threshold
|
dict
|
Threshold values for temporal consistency by index. Example: {"Ndvi": 0.06, "Lai": 0.3, "S2Rep": 2.5} If None and compute_temporal_consistency=True, uses default thresholds. |
None
|
mode
|
str
|
Output mode. Default: 'full' - 'full': Returns both raw and smoothed data merged on date - 'raw': Returns only raw observation data (no smoothed interpolation) |
'full'
|
partial_frequency
|
int
|
Partial save frequency during bulk processing. Default: 50 |
50
|
kpi_filter
|
dict
|
KPI configuration for aggregating time series data: - 'kpi_name' (str): Name of the KPI - 'aggregation' (str): Type of aggregation (accumulation, average, max, min, count_gt, count_lt, count_between, std) - 'threshold' (float or tuple, optional): Threshold value(s) for count operations - Single float for count_gt/count_lt - Tuple of two floats (min, max) for count_between |
None
|
get_mrts_api 🔗
Get MRTS (Medium Resolution Time Series) data from API using geometry.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data containing: - 'geometry' (str): WKT geometry string - 'id' (str, optional): Season field ID - 'crop' (str, optional): Crop type - 'sowing_date' (str, optional): Sowing date in YYYY-MM-DD format - 'start_date' (str, optional): Start date in YYYY-MM-DD format - 'end_date' (str, optional): End date in YYYY-MM-DD format If start_date/end_date not provided, uses dates from mrts_params |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
JSON response with MRTS data |
get_mrts_api_safe 🔗
Safe wrapper around get_mrts_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain 'geometry' and optionally 'id', 'start_date', 'end_date' |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str |
dict
|
} |
format_mrts_json 🔗
Process MRTS API response into a DataFrame with date, raw and smoothed values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_mrts_json
|
dict
|
JSON response with 'rawData' and 'smoothedData' keys |
required |
entity_data
|
dict
|
Original entity data to include ID in output |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: - mode='full': entity_id (optional), date, raw_value, noised, temporalConsistencyCheck, mask, coveragePercent, image_id, smoothed_value - mode='raw': entity_id (optional), date, raw_value, noised, temporalConsistencyCheck, mask, coveragePercent, image_id |
process_single_entity_mrts 🔗
Process a single entity for medium resolution time series extraction with optional KPI filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
pd.DataFrame, pd.Series, or dict
|
Entity data with required fields: - 'id': Entity identifier - 'geometry': WKT geometry string - 'start_date' (optional): Start date (YYYY-MM-DD) - falls back to params - 'end_date' (optional): End date (YYYY-MM-DD) - falls back to params - 'years' (optional): List of historical years for KPI comparison If DataFrame with multiple rows: only first row is processed |
required |
params
|
dict
|
Override mrts_params (includes kpi_filter config and fallback dates) |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains 'data' (DataFrame) and 'error' - If KPI requested: DataFrame with metadata + KPI values (one row per entity) - If NO KPI: DataFrame with metadata + daily date/value data (multiple rows) |
process_mrts_bulk_parallel 🔗
process_mrts_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='mrts', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity medium resolution vegetation time series using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain 'id' and 'geometry'. |
required |
params
|
dict
|
Vegetation time series parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "mrts". |
'mrts'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts. |
WeatherExtractor 🔗
Bases: BaseExtractor
Extracts field-level weather data for agricultural entities.
Retrieves historical daily weather data and agro-meteorological indices for a given geometry and date range. Supports configurable weather parameters and KPI aggregation.
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_weather.ipynb
Args (setup_weather_parameters): weather_type (str): Weather data type ('HISTORICAL_DAILY'). Default: 'HISTORICAL_DAILY' weather_parameters (str|list): Parameters to retrieve; 'none' expands to all available parameters. Default: 'precipitation.cumulative' kpi_filter (dict): KPI aggregation config (kpi_name, aggregation, threshold) historical_years (int): Number of prior years of weather history to fetch. Default: 0 page_limit (int): Override the API row limit ($limit); by default it is auto-sized to the query span so multi-year ranges are not truncated. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); start_date, end_date (optional overrides)
Output columns
entity_id, date, + dynamic weather parameter columns (temperature, precipitation, etc.)
get_new_token 🔗
Implements token refresh logic for WeatherExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_weather_parameters 🔗
setup_weather_parameters(weather_type='HISTORICAL_DAILY', weather_parameters='precipitation.cumulative', historical_years=0, partial_frequency=50, kpi_filter=None, page_limit=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure params for weather extraction for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
weather_type
|
str
|
Type of weather data (HISTORICAL_DAILY, FORECAST_DAILY, FORECAST_HOURLY) |
'HISTORICAL_DAILY'
|
weather_parameters
|
str or list
|
Weather parameters to extract. Defaults to 'precipitation.cumulative'. Pass 'none' to retrieve every available parameter. |
'precipitation.cumulative'
|
historical_years
|
int or list
|
Number of historical years for KPI comparison, or list of specific years |
0
|
partial_frequency
|
int
|
How often to save partial results |
50
|
kpi_filter
|
dict
|
KPI configuration |
None
|
page_limit
|
int
|
Override the API row limit ($limit). By default the limit is auto-sized to the query span (see get_weather_data) so long multi-year ranges are not truncated; set this to force an explicit value. Must be a positive integer. Default: None (auto-size). |
None
|
get_weather_data 🔗
Get weather data for selected parameters using self.weather_params
The API returns one row per day in ascending date order and caps the response at
$limit rows, so a fixed limit silently truncates long ranges to the earliest N
days (dropping the most recent data — the current-period window for KPIs). The row
limit is therefore auto-sized to the query span via
api_utils.autosize_row_limit (floor 1000 keeps short single-season/forecast
requests byte-identical). A page_limit set in setup_weather_parameters wins
over the auto-sized value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data with geometry, start_date, end_date |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Weather API response |
get_weather_data_safe 🔗
Safe wrapper around get_weather_data(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain 'id', 'geometry', 'start_date', 'end_date' keys |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str |
dict
|
} |
format_weather_json 🔗
Format weather API response into a DataFrame. Dynamically handles all weather parameters returned by the API.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_weather_json
|
list
|
JSON response from get_weather_data |
required |
entity_data
|
dict
|
Original entity data to include ID in output |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: DataFrame with date and all weather parameter columns |
process_single_entity_weather 🔗
Process a single entity for weather data extraction with optional KPI filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
pd.DataFrame, pd.Series, or dict
|
Entity data with required fields |
required |
params
|
dict
|
Override weather_params |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains 'data' (DataFrame) and 'error' |
process_entity_weather_bulk_parallel 🔗
process_entity_weather_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='weather', generate_report=False, report_options=None, use_cache=None, spatial_grouping=False, spatial_precision=DEFAULT_SPATIAL_PRECISION, spatial_max_window_days=DEFAULT_MAX_WINDOW_DAYS)
Bulk processing of entity weather data using threads + progress bar, with optional fail-safe retry and filter capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process |
required |
params
|
dict
|
Weather parameters override |
None
|
max_workers
|
int
|
Number of threads to use |
5
|
output_path
|
str
|
Directory to save results |
None
|
partial_frequency
|
int
|
How often to save partial results |
50
|
fail_safe
|
bool
|
If True, retries only previously failed entities |
False
|
filter_column
|
str
|
Column name to filter entities by |
None
|
filter_value
|
any
|
Value to filter |
None
|
filter_type
|
str
|
'exclude' or 'include' |
'exclude'
|
merge_existing
|
str
|
Merge strategy |
None
|
skip_export
|
bool
|
If True, skip final export |
False
|
prefix
|
str
|
Prefix for output filenames |
'weather'
|
use_cache
|
bool
|
Enable/disable caching for this call |
None
|
spatial_grouping
|
bool
|
If True, group fields by |
False
|
spatial_precision
|
int
|
Geohash precision for the bucket (default
|
DEFAULT_SPATIAL_PRECISION
|
spatial_max_window_days
|
int
|
Cap on a group's widened window (default
|
DEFAULT_MAX_WINDOW_DAYS
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics |
cropidExtractor 🔗
Bases: BaseExtractor
Extracts crop identification analytics for agricultural entities.
AI-powered crop type classification using multi-temporal satellite imagery. Supports end-of-season and in-season mask types, historical crop rotation analysis, and configurable confidence thresholds.
Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_cropid.ipynb
Args (setup_cropid_parameters): begin_year (int): Start year for historical analysis. Default: 2020 end_year (int): End year for analysis. Default: 2025 mask_type (str): Mask type ('EndSeason', 'InSeason', 'PreSeason'). Default: 'EndSeason' limit_nb_crop (int): Max number of crop candidates per field. Default: 1 crop_mask_percent (int): Minimum crop mask percentage. Default: 50 mode (str): Output mode — 'history' (one row per (year, crop)), 'year' (filter to entity crop, long-form), 'historical_season' (comma-separated string of matching years), 'full_history' (wide-form, one row per entity, one column per year in begin_year..end_year filled with the top crop code per year). Default: 'history' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry (required); crop (required for 'year' and 'historical_season' modes)
Output columns
Depend on mode:
- 'history' : entity_id, year, eda_crop_code, crop_name, raw_crop_name, crop_mask_percent
- 'year' : entity_id, year, eda_crop_code
- 'historical_season' : entity_id, historical_season (comma-separated years)
- 'full_history' : entity_id + one column per year in [begin_year, end_year],
filled with the top crop code (highest cropMaskPercent) per year
get_new_token 🔗
Implements token refresh logic for cropidExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.
setup_cropid_parameters 🔗
setup_cropid_parameters(begin_year=2020, end_year=2025, mask_type='EndSeason', limit_nb_crop=1, crop_mask_percent=50, mode='history', partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for Crop ID extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
begin_year
|
int
|
Starting year (YYYY format, e.g., 2020) |
2020
|
end_year
|
int
|
Ending year (YYYY format, e.g., 2025) |
2025
|
mask_type
|
str
|
Type of crop mask. Options: 'InSeason', 'EndSeason', 'PreSeason' |
'EndSeason'
|
limit_nb_crop
|
int
|
Maximum number of crops to return per field |
1
|
crop_mask_percent
|
int
|
Minimum percentage of field covered by crop (0-100) |
50
|
mode
|
str
|
Processing mode — one of: - 'history' : all crops/years (long-form, one row per (year, crop)) - 'year' : filtered to entity crop (long-form) - 'historical_season' : comma-separated string of matching years - 'full_history' : wide-form pivot, one row per entity with one column per year (begin_year..end_year) filled with the top crop code per year |
'history'
|
partial_frequency
|
int
|
How often to save partial results during bulk processing |
50
|
Returns:
| Type | Description |
|---|---|
|
None |
get_cropid_api 🔗
Request crop id for an entity
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
entity data (id, geometry, crop) |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
JSON response from the crop ID API |
get_cropid_api_safe 🔗
Safe wrapper around get_cropid_api(). Returns structured response with success flag, data or error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Must contain a 'geometry' key in WKT format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str |
dict
|
} |
format_cropid_json 🔗
Format crop ID API response into a DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
JSON response from get_cropid_api |
required |
entity_data
|
dict
|
Original entity data to include ID in output |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: DataFrame with columns: entity_id (optional), year, crop_name, raw_crop_name, crop_mask_percent, eda_crop_code |
format_cropid_year_json 🔗
Format crop ID API response and filter to only years matching the entity's crop.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
JSON response from get_cropid_api |
required |
entity_data
|
dict
|
Original entity data with 'crop' as string (e.g., {'crop': 'SOYBEANS'}) |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: DataFrame with entity_id, year, eda_crop_code for matching crops only |
format_cropid_historical_season_json 🔗
Format crop ID API response to return a comma-separated string of years where the crop matches the entity's crop.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
JSON response from get_cropid_api |
required |
entity_data
|
dict
|
Entity data with 'crop' as string (e.g., {'crop': 'SOYBEANS'}) |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Comma-separated years matching the entity crop (e.g., "2020,2022,2024,2025") |
format_cropid_full_history_json 🔗
Format crop ID API response into a wide-form, single-row DataFrame:
one column per year in [begin_year, end_year], filled with the top
crop code (highest cropMaskPercent) for that year — or NaN if
the year is absent from the response.
Year columns are string-typed (e.g. "2020") so the wide frame plays
nicely with parquet and bulk concat across entities. When multiple crops
are reported for one year (limit_nb_crop > 1 upstream), only the top
crop by cropMaskPercent is kept; tie-break is the API's natural
ordering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
JSON response from get_cropid_api |
required |
entity_data
|
dict
|
Entity data — must include 'id' |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: 1-row DataFrame with columns |
process_single_entity_cropid 🔗
Process a single entity for crop ID extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict
|
Entity data with 'id', 'geometry', and 'crop' |
required |
params
|
dict
|
Override cropid_params |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains 'data' (DataFrame) and 'error' (dict or None) |
process_cropid_bulk_extraction_parallel 🔗
process_cropid_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='cropid', generate_report=False, report_options=None, use_cache=None)
Bulk processing of cropid requests using threads + progress bar, with optional fail-safe retry, partial save capabilities, and caching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
Entities to process (must contain 'id', 'geometry', and 'crop'). |
required |
params
|
dict
|
Override monitoring parameters. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final results and error logs. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export (useful when chaining extractions). Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "cropid". |
'cropid'
|
use_cache
|
bool
|
Override instance-level cache setting. Default: None (uses self.use_cache). |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] } |
LocationBasedBorderExtractor 🔗
Bases: BaseExtractor
Extracts a field-border polygon from a single longitude/latitude location.
Calls the Geosys /field-borders/v1/AutomaticBoundary endpoint with a
location=lon,lat query parameter and returns the polygon (in WKT) of
the field that contains that point. Useful for converting a list of
geocoded locations (CSV / API output) into proper field geometries.
Documentation: https://api.geosys-na.net/field-borders/v1/swagger/index.html Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_LocationBasedBorder.ipynb
Args (setup_location_based_border_parameters):
simplified_geom (bool): If True, returns a simplified field geometry
(lower shape-point count). Default: True. Maps to the API's
simplified_geom query parameter.
partial_frequency (int): How often to flush partial results. Default: 50.
use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping):
id (required) - used to stamp results, log lines, and cache keys.
geometry (required) - must be a Point WKT (e.g. POINT(-93.6 41.5)).
If the row carries a non-point geometry the entity is rejected with
a clean error rather than being silently degraded to a centroid.
Output columns
entity_id, point_geometry (echo of input point as WKT),
polygon_geometry (returned field border as WKT),
plus any flat scalar fields the API returns alongside geometry
(e.g. area, sourceId).
setup_location_based_border_parameters 🔗
setup_location_based_border_parameters(simplified_geom: bool = True, partial_frequency: int = 50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for location-based-border extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
simplified_geom
|
bool
|
If True, returns simplified field geometry. Default: True. |
True
|
partial_frequency
|
int
|
How often to flush partial results. Default: 50. |
50
|
column_mapping
|
dict
|
Column name overrides (e.g. |
None
|
output_mapping
|
dict
|
Output column renames. |
None
|
exclude_columns
|
list
|
Columns to exclude from output. |
None
|
output_columns
|
list
|
Whitelist of output columns. |
None
|
use_cache
|
bool
|
Whether to use caching for this run. |
None
|
get_location_based_border_api 🔗
Call the AutomaticBoundary API for a single entity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict | Series
|
Entity data with |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Parsed JSON response from the API. Always carries the input |
|
|
|
||
|
downstream formatter can stamp them. |
get_location_based_border_api_safe 🔗
Safe wrapper around get_location_based_border_api().
Returns:
| Name | Type | Description |
|---|---|---|
dict |
|
format_location_based_border_json 🔗
Normalize an AutomaticBoundary response into a single-row DataFrame.
Accepts three shapes (the API's response schema isn't documented in the Swagger so we tolerate the common variants):
- Flat dict with a
geometryfield — string WKT or GeoJSON geometry. - GeoJSON Feature:
{"type": "Feature", "geometry": {...}, "properties": {...}}. - GeoJSON FeatureCollection — uses
features[0].
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
API response (augmented with |
required |
params
|
dict
|
Override params (unused here, kept for symmetry). |
None
|
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Single-row DataFrame with |
|
|
|
|
|
properties returned by the API. |
process_single_entity_location_based_border 🔗
Process location-based-border extraction for a single entity with retry logic.
Returns:
| Name | Type | Description |
|---|---|---|
dict |
|
process_location_based_border_bulk_extraction_parallel 🔗
process_location_based_border_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='location_based_border', generate_report=False, report_options=None, use_cache=None)
Bulk processing of location-based-border requests using threads + progress bar.
Args mirror the family pattern (coverage_function / difference_functions /
processor_change_index_functions). Returns the standard
{results_df, global_errors, total_entities, total_calculations,
successful_calculations, failed_calculations, failed_ids} dict.
ZoningExtractor 🔗
Bases: BaseExtractor
Extracts management zone maps (SAMZ) for agricultural entities.
Calls the SAMZ (SAtellite derived Management Zones) API to get management zones from satellite imagery. Returns field-level statistics including variability, productivity indices, and per-zone area/productivity breakdowns.
Documentation: https://docs.earthdaily.com/agro/library/Field%20Level%20Maps/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_zoning.ipynb
Args (setup_zoning_parameters): num_zones (int): Number of management zones to generate. Default: 5 output_epsg (int): Output coordinate system. Default: 4326 postprocess (str): Processing mode - 'stats', 'stats_geo', 'links', or 'file'. Default: 'stats' 'stats' returns one field-level row; 'stats_geo' returns one row per zone, carrying that zone's geometry alongside the field-level stats. map_format (str): Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None output_path (str): Directory for file downloads. Required when postprocess='file'. directLinks (bool): Request direct download links from API. Default: False partial_frequency (int): How often to save partial results. Default: 50 use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None
Entity fields (via column_mapping): id, geometry, image_id (required)
Output columns
Varies by postprocess mode. stats: entity_id, field_variability, most_variable_zone, highest_interzone_variability, field_productivity_index, field_variability_index, zone_{n}area_percent, zoneproductivity_index, zonevariability_index stats_geo: one row per zone — zone_name, zone_geometry, productivity_index, variability_index, area_percent, number_of_pixels, zone_mean, zone_max, zone_min, zone_area, plus the field-level stats above repeated on every row links: image_png_link, worldfile_link, thumbnail_link, worldfile, bbox_, map_width, map_height file: status, map_format, file_count, total_size_bytes, saved_files
TODO
- Cache support: cache is disabled by default because results depend on image_id (which can be a list). When image_id is a list, it gets stored as a list object in the DataFrame, which does not deduplicate cleanly in parquet. To enable cache properly, image_id lists need to be serialized to a stable string (e.g. sorted, joined with ";") before cache lookup and storage, so that the same set of images always produces the same cache key regardless of order.
setup_zoning_parameters 🔗
setup_zoning_parameters(num_zones=5, output_epsg=4326, postprocess='stats', map_format=None, output_path=None, skip_existing=True, directLinks=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)
Configure parameters for management zone extraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_zones
|
int
|
Number of management zones (2-10). Default: 5 |
5
|
output_epsg
|
int
|
Output EPSG code for projection. Default: 4326 |
4326
|
postprocess
|
str
|
Processing mode - 'stats', 'stats_geo', 'links', or 'file'. Default: 'stats' 'stats' returns a single field-level row per entity. 'stats_geo' returns one row per zone, adding the zone geometry (segments merged into a GEOMETRYCOLLECTION when a zone has several) and its per-zone mean/max/min/area. |
'stats'
|
map_format
|
str
|
Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None |
None
|
output_path
|
str
|
Directory for file downloads. Required when postprocess='file'. |
None
|
skip_existing
|
bool
|
Skip download if file already exists. Default: True |
True
|
directLinks
|
bool
|
Request direct download links. Default: False Note: Automatically set to True when postprocess='links' |
False
|
partial_frequency
|
int
|
How often to save partial results. Default: 50 |
50
|
column_mapping
|
dict
|
Column name overrides. |
None
|
output_mapping
|
dict
|
Output column renames. |
None
|
exclude_columns
|
list
|
Columns to exclude from output. |
None
|
output_columns
|
list
|
Whitelist of output columns. |
None
|
use_cache
|
bool
|
Whether to use caching. |
None
|
get_zoning_map_api 🔗
Call the SAMZ API to generate management zones for an entity. Returns raw Response when map_format is set (file download), or parsed JSON otherwise (stats/links).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id', 'geometry', and 'image_id' image_id can be a single string or a list of image ID strings. |
required |
Returns:
| Type | Description |
|---|---|
|
requests.Response (file mode) or dict (stats/links mode) |
get_zoning_map_api_safe 🔗
Safe wrapper around get_zoning_map_api().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id', 'geometry', and 'image_id' |
required |
Returns:
| Name | Type | Description |
|---|---|---|
dict |
dict
|
{"success": bool, "data": dict|None, "error": str|None, "seasonfield_id": str} |
format_zoning_stats_json 🔗
Extract field-level and per-zone statistics from SAMZ API response.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
Parsed API response |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Single-row DataFrame with field stats and per-zone columns |
format_zoning_stats_geo_json 🔗
Extract per-zone statistics with geometry, productivity index, and variability index.
Produces one row per zone, combining data from legend.ranges (productivity/variability) and zones (geometry/stats). Zone geometries from multiple segments are merged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict
|
Parsed API response |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: One row per zone with geometry, indices, and field-level stats |
format_zoning_links_json 🔗
Extract direct links and metadata from SAMZ API response.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
response_json
|
dict or Response
|
API response with _links, worldFile, etc. |
required |
Returns:
| Type | Description |
|---|---|
|
pd.DataFrame: Single-row DataFrame with links and spatial metadata |
format_zoning_file 🔗
Save SAMZ zone map file response (PNG, TIFF.ZIP, or SHP.ZIP) to disk.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_data
|
dict
|
Entity data including 'id' |
required |
image_id
|
str or list
|
Image ID(s) for filename. If a list, joined with '_' for the filename. |
required |
response
|
Response
|
Raw response from get_zoning_map |
required |
output_path
|
str
|
Directory to save files. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Status and file path information |
process_single_entity_zoning 🔗
Process zoning extraction for a single entity with retry logic. Routes to stats, links, or file processing based on postprocess parameter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
dict | Series
|
Entity data with id, geometry, image_id |
required |
params
|
dict
|
Override parameters |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
{"data": DataFrame or None, "error": dict or None} |
process_entity_zoning_bulk_parallel 🔗
process_entity_zoning_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='zoning', generate_report=False, report_options=None, use_cache=None)
Bulk processing of entity SAMZ zones using threads + progress bar, with optional fail-safe retry and filter capabilities.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
entity_list
|
DataFrame
|
List of entities to process, must contain 'id', 'geometry', and 'image_id'. |
required |
params
|
dict
|
Zoning parameters override. |
None
|
max_workers
|
int
|
Number of threads to use. |
5
|
output_path
|
str
|
Directory to save final CSV results. |
None
|
partial_frequency
|
int
|
How often to save partial results. |
50
|
fail_safe
|
bool
|
If True, per-entity failures are captured to
|
False
|
filter_column
|
str
|
Column name to filter entities by. |
None
|
filter_value
|
any
|
Value to filter in the filter_column. |
None
|
filter_type
|
str
|
'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'. |
'exclude'
|
merge_existing
|
str
|
Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default. |
None
|
skip_export
|
bool
|
If True, skip final export. Default: False. |
False
|
prefix
|
str
|
Prefix for output filenames and failed IDs files. Default: "zoning". |
'zoning'
|
generate_report
|
bool
|
Generate HTML report. Default: False. |
False
|
report_options
|
dict
|
Report configuration options. |
None
|
use_cache
|
bool
|
Enable/disable caching for this call. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict |
Contains results DataFrame, errors list, and summary statistics. |