Skip to content

Foundational extractors🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

CoverageExtractor 🔗

Bases: BaseExtractor

Extracts satellite image coverage analytics for agricultural entities.

Retrieves available satellite imagery metadata (coverage percentage, sensor, date, mask) for a given geometry and date range. Supports duplicate filtering across sensors and multi-sensor coverage analysis.

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_coverage.ipynb

Args (setup_coverage_parameters): vegetation_index (str): Vegetation index type ('NDVI', 'EVI', etc.). Default: 'NDVI' start_date (str): Start date in YYYY-MM-DD format. Default: '2025-01-01' end_date (str): End date in YYYY-MM-DD format. Default: None — no upper bound, i.e. every image from start_date onward is returned (the query becomes image.date=$gte:start_date). It is not defaulted to today; set it explicitly to bound the catalogue query. clear_cover_min (int): Minimum clear sky coverage percentage (0-100). Default: 90 clear_cover_max (int): Maximum clear sky coverage percentage (0-100). Default: 100 use_specific_date (bool): Restrict to exact date. Default: False filter (str): Filtering mode - 'none', 'duplicate' or 'crop_coverage'. Default: 'none' 'duplicate' collapses near-simultaneous multi-sensor captures (see delay). 'crop_coverage' expands the query per entity across prior seasons and requires historical_seasons; the per-entity window is computed from the entity's start_date/end_date month-day applied to each historical year. delay (int): Max days between multi-sensor captures for duplicate filter. Default: 3 mask (str): Cloud mask type — 'Auto', 'Native', 'ACM', 'ML', 'MLCirrus' or 'All' (case-insensitive). 'All' returns one row per image per mask, so the same image appears several times and the 'mask' output column tells them apart — use it to compare masks. Not compatible with filter='duplicate'. Default: 'auto' recalibration (bool): Apply cross-sensor recalibration to coverage metrics. Default: False historical_seasons (list[int]): Prior season years used when filter='crop_coverage' (required for that filter, e.g. [2024, 2023, 2022]). Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required)

Output columns

image_id, coverage_percent, date, mask, sensor, spatial_resolution

get_new_token 🔗

get_new_token()

Implements token refresh logic for CoverageExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_coverage_parameters 🔗

setup_coverage_parameters(vegetation_index='NDVI', start_date='2025-01-01', end_date=None, clear_cover_min=90, clear_cover_max=100, use_specific_date=False, filter='none', delay=3, mask='auto', partial_frequency=50, recalibration=False, historical_seasons=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure params for imagery catalog extraction for an entity.

When filter='crop_coverage', historical_seasons must be provided. The extraction period is computed per entity from start_date/end_date fields across all historical years (smallest year start to largest year end).

get_satellite_coverage_by_geometry 🔗

get_satellite_coverage_by_geometry(entity_data: dict)

Query the coverage API to get imagery metadata for a given entity geometry. If recalibration is enabled in coverage_params, uses the recalibration endpoint and includes crop in the payload.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format. If recalibration=True, must also contain a 'crop' key.

required

Returns:

Name Type Description
dict

API response with imagery metadata

get_satellite_coverage_by_geometry_safe 🔗

get_satellite_coverage_by_geometry_safe(entity_data: dict) -> dict

Safe wrapper around get_satellite_coverage_by_geometry(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_coverage_json 🔗

format_coverage_json(response_coverage_json, entity_id='unknown')

Process API response to extract image ID, coverage percent, formatted date, sensor, spatial resolution, and mask information.

Parameters:

Name Type Description Default
response_coverage_json str | list | dict

API response (stringified JSON, parsed list, or dict).

required
entity_id str

Entity ID for logging purposes.

'unknown'

Returns:

Type Description

pd.DataFrame: Cleaned DataFrame of imagery coverage.

process_single_entity_coverage 🔗

process_single_entity_coverage(row, params=None)

Process coverage extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict

Entity data containing id and geometry

required
params dict

Coverage parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_entity_coverage_bulk_parallel 🔗

process_entity_coverage_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='coverage', flm_params=None, recalibration=False, generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity coverage using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain 'id' and 'geometry'.

required
params dict

Coverage parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "coverage".

'coverage'
flm_params dict

FLM parameters. If provided, sets recalibration=True and stores on self.

None
recalibration bool

If True, uses recalibration API endpoint. Default: False.

False
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.

FLMExtractor 🔗

Bases: BaseExtractor

Extracts Field Level Maps (vegetation index maps) for agricultural entities.

Generates field-level vegetation index maps (NDVI, EVI, NDRE, etc.) from satellite imagery. Supports map statistics extraction, direct download links, and raster file export with configurable clipping and buffering.

Documentation: https://docs.earthdaily.com/agro/library/Field%20Level%20Maps/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_FLM_extraction_functions.ipynb

Args (setup_flm_parameters): vegetation_index (str): Index type ('NDVI', 'EVI', 'NDRE', etc.). Default: 'NDVI' map_format (str): Output format (None, 'png', 'tiff'). Default: None output_epsg (int): Output coordinate system. Default: 4326 postprocess (str): Post-processing mode ('stats', 'links', 'file', 'histogram'). Default: 'stats' 'histogram' sets histogram=true on the query and returns the per-bucket breakdown as well as the global min/max/mean. extract_stats (bool): Extract map statistics. Default: False directLinks (bool): Return direct download links. Default: False clipping (str): Clipping mode ('FieldBorder', 'None'). Default: 'FieldBorder' buffer (int): Buffer around field border in meters. Default: 0 number_bins (int): Number of legend bins for the rendered map. Default: None legendType (str): Legend style for the rendered map. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required)

Output columns

Varies by postprocess mode: stats: stat_max, stat_mean, stat_min links: image_png_link, worldfile_link, thumbnail_link, bbox_*, map_width, map_height file: status, map_format, file_count, total_size_bytes, saved_files histogram: stat_min, stat_max, stat_mean, value_min_{i}, value_max_{i}, num_pixels_{i}, area_{i} (one set of {i} columns per bucket, single row per entity)

get_new_token 🔗

get_new_token()

Implements token refresh logic for FLMExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_flm_parameters 🔗

setup_flm_parameters(vegetation_index='NDVI', map_format=None, output_epsg=4326, postprocess='stats', skip_existing=True, output_path=None, extract_stats=False, partial_frequency=50, directLinks=False, clipping='FieldBorder', buffer=0, number_bins=None, legendType=None, use_cache=None)

Configure parameters for FLM extraction

Parameters:

Name Type Description Default
vegetation_index str

Vegetation index to extract

'NDVI'
map_format str

Format of the map ('png', 'tiff.zip', 'shp.zip')

None
output_epsg int

Output EPSG code for projection

4326
postprocess str

Post-processing type - 'stats', 'file', 'links', or 'histogram'. Default: 'stats'

'stats'
skip_existing bool

Skip download if file already exists

True
output_path str

Path to save downloaded maps

None
extract_stats bool

If True, extract statistics only (no file download)

False
partial_frequency int

How often to save partial results

50
directLinks bool

Whether to use direct links in API response. Default: False Note: Automatically set to True when postprocess='links'

False
clipping str

Clipping method - 'FieldBorder' or 'Bbox'. Default: 'FieldBorder' Note: Only applicable for COLORCOMPOSITION vegetation index

'FieldBorder'
buffer int

Buffer distance in meters around geometry. Default: 0

0
number_bins int

Number of histogram bins (1-255). When set, appends &numberOfBins=<value> to the API URL. Default: None (API uses its built-in default).

None
legendType str

Legend type - one of 'Fixed', 'Dynamic', 'Common'. When set, appends &legendType=<value> to the API URL. Default: None (API uses 'Dynamic').

None

get_flm_map 🔗

get_flm_map(entity_data: dict, image_id, file_extension=None)

Process FLM map: either download file OR extract statistics OR get direct links Uses parameters from self.flm_params (set by setup_parameters)

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id' and 'geometry'

required
image_id str

Image ID to download

required
file_extension str

File extension override

None

Returns:

Type Description

requests.Response: Raw API response

get_flm_map_safe 🔗

get_flm_map_safe(entity_data: dict, image_id) -> dict

Safe wrapper around get_flm_map(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required
image_id str

Image ID

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_flm_map_stats_json 🔗

format_flm_map_stats_json(response_flm_map, range=False)

Extract legend statistics and metadata from Map Reference JSON. Converts raw response to JSON at the beginning.

Parameters:

Name Type Description Default
response_flm_map dict or Response

API response JSON with legend, mapSize, bBox, etc.

required
range bool

If True, also return ranges DataFrame. Default is False.

False

Returns:

Type Description

pd.DataFrame or tuple: - If range=False: stats_df only - If range=True: (stats_df, ranges_df)

format_flm_map_histogram_json 🔗

format_flm_map_histogram_json(response_flm_map)

Extract histogram stats and per-range values from Map Reference JSON.

Produces a single-row DataFrame per entity: - global stats: stat_min, stat_max, stat_mean - per-bucket columns (1-indexed): value_min_{i}, value_max_{i}, num_pixels_{i}, area_{i}

Parameters:

Name Type Description Default
response_flm_map dict or Response

API response carrying a histogram block (requires histogram=true on the query, which is set automatically when postprocess='histogram').

required

Returns:

Type Description

pd.DataFrame: single-row wide DataFrame.

format_flm_map_file 🔗

format_flm_map_file(entity_data: dict, image_id: str, response, output_path=None)

Save FLM map file response (PNG, TIFF.ZIP, or SHP.ZIP) to disk. Handles extraction of zipped files and proper naming.

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id' and optional 'name'

required
image_id str

Image ID for filename

required
response Response

Raw response from get_flm_map

required
output_path str

Directory to save files. Uses self.flm_params['output_path'] if not provided.

None

Returns:

Name Type Description
dict

Status and file path information

format_flm_links_json(response_flm_map)

Extract direct links and metadata from FLM API response. Converts raw response to JSON at the beginning.

Parameters:

Name Type Description Default
response_flm_map dict or Response

API response JSON with _links, worldFile, etc.

required

Returns:

Type Description

pd.DataFrame: DataFrame containing links and metadata (one row per entity)

process_single_entity_flm 🔗

process_single_entity_flm(row, params=None)

Process FLM extraction for a single entity with retry logic. Supports four postprocess modes: 'stats', 'links', 'file', or 'histogram'

Parameters:

Name Type Description Default
row dict

Entity data containing id, geometry, and image_id

required
params dict

FLM parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_entity_flm_bulk_parallel 🔗

process_entity_flm_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='flm', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity FLM maps using threads + progress bar, with optional fail-safe retry and filter capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain 'id', 'geometry', and 'image_id'.

required
params dict

FLM parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "flm".

'flm'
use_cache bool

Enable/disable caching for this call.

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics.

VegationTsExtractor 🔗

Bases: BaseExtractor

Extracts high-resolution vegetation time series for agricultural entities.

Retrieves historical and near-real-time vegetation index time series (NDVI, EVI, LAI, etc.) at field level. Supports period-based and target-date extraction modes, KPI aggregation, and multi-year historical comparison.

Documentation: https://docs.earthdaily.com/agro/library/Vegetation_time_series/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_VTS.ipynb

Args (setup_vegetation_ts_parameters): start_date (str): Start date in YYYY-MM-DD format end_date (str): End date in YYYY-MM-DD format vegetation_index (str): Index type ('NDVI', 'EVI', 'CVI', 'GNDVI', 'NDWI', 'LAI'). Default: 'NDVI' is_extrapolated (bool): Include extrapolated values. Default: True limit (int): Max number of records. Default: 3000 historical_years (int): Number of historical years. Default: 10 extraction_mode (str): 'period' or 'target_dates'. Default: 'period' target_dates (list): Specific dates to extract (for target_dates mode) kpi_filter (dict): KPI aggregation config (kpi_name, aggregation, threshold) use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop (required for LAI); start_date, end_date (optional overrides)

Output columns

entity_id, date, value (vegetation index), + historical year columns when applicable

get_new_token 🔗

get_new_token()

Implements token refresh logic for VegationTsExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_vegetation_ts_parameters 🔗

setup_vegetation_ts_parameters(start_date: str = None, end_date: str = None, vegetation_index='NDVI', is_extrapolated=True, limit=3000, historical_years=10, partial_frequency=50, extraction_mode='period', target_dates=None, kpi_filter=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for vegetation index extraction.

Parameters:

Name Type Description Default
start_date str

Baseline start date in YYYY-MM-DD format. Used as default when entity_data does not provide its own start_date.

None
end_date str

Baseline end date in YYYY-MM-DD format. Used as default when entity_data does not provide its own end_date.

None
vegetation_index str

Vegetation index to extract (NDVI, EVI)

'NDVI'
is_extrapolated bool

Whether to include extrapolated values

True
limit int

Maximum number of records to retrieve

3000
historical_years int

Number of historical years to include (max 15)

10
partial_frequency int

How often to save partial results

50
extraction_mode str

Extraction mode. - 'period': seasonal MM-DD window extraction. Uses start_date/end_date as baseline (entity_data can override). Expands by historical_years for KPI comparison. - 'windows': absolute date range extraction. Uses start_date/end_date as baseline (entity_data can override). No historical expansion. - 'specific_dates': punctual extraction at exact target dates.

'period'
target_dates list
  • List of specific dates (YYYY-MM-DD) when mode='specific_dates'
  • Not used for 'period' or 'windows' modes
None
kpi_filter dict

KPI configuration

None

process_single_entity_specific_dates 🔗

process_single_entity_specific_dates(row, params=None)

Process a single entity for vegetation time series extraction at specific dates.

Parameters:

Name Type Description Default
row Series or dict

Entity data with 'id' and 'name'

required
params dict

Override vegetation_ts_params

None

Returns:

Name Type Description
dict

Contains 'data' (DataFrame with columns: date, field_name, ndvi_value) and 'error'

get_vegetation_api 🔗

get_vegetation_api(entity_data: dict)

Get vegetation time series data from VTS API.

Parameters:

Name Type Description Default
entity_data dict

Entity data containing: - 'id' (str): Season field ID - 'start_date' (str): Start date in YYYY-MM-DD format (required for 'period' mode) - 'end_date' (str): End date in YYYY-MM-DD format (required for 'period' mode) Note: For 'windows' mode, start_date and end_date are optional and will use target_dates from vegetation_ts_params if not provided

required

Returns:

Name Type Description
dict

JSON response with vegetation time series data

get_vegetation_api_safe 🔗

get_vegetation_api_safe(entity_data: dict) -> dict

Safe wrapper around get_vegetation_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain 'id', 'start_date', and 'end_date' keys

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str

dict

}

format_vegetation_ts_json 🔗

format_vegetation_ts_json(response_vegetation_ts_json, entity_data: dict = None)

Process vegetation API response into a DataFrame with date and value.

Parameters:

Name Type Description Default
response_vegetation_ts_json list

JSON response from get_vegetation_api (list of date/value records)

required
entity_data dict

Original entity data to include ID in output

None

Returns:

Type Description

pd.DataFrame: DataFrame with columns: entity_id (optional), date, value

process_single_entity_vegetation_ts 🔗

process_single_entity_vegetation_ts(row, params=None)

Process a single entity for vegetation time series extraction with optional KPI filtering.

Parameters:

Name Type Description Default
row Series or dict

Entity data with required fields: - 'id': Entity identifier - 'geometry': WKT geometry string - 'start_date': Start date (YYYY-MM-DD) - 'end_date': End date (YYYY-MM-DD) - 'years' (optional): List of historical years for KPI comparison

required
params dict

Override vegetation_ts_params (includes kpi_filter config)

None

Returns:

Name Type Description
dict

Contains 'data' (DataFrame) and 'error' - If KPI requested: DataFrame with metadata + KPI values (one row per entity) - If NO KPI: DataFrame with metadata + daily date/value data (multiple rows)

process_entity_vegetation_ts_bulk_parallel 🔗

process_entity_vegetation_ts_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='vegetation_ts', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity vegetation time series using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain 'id' and 'geometry'.

required
params dict

Vegetation time series parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "vegetation_ts".

'vegetation_ts'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.

process_specific_dates_bulk_parallel 🔗

process_specific_dates_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', skip_export=False, prefix='vegetation_specific_dates', generate_report=False, report_options=None, use_cache=None)

Bulk processing of vegetation time series at specific dates using parallel threads, with optional caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities with 'id' and 'name' columns

required
params dict

Override vegetation_ts_params

None
max_workers int

Number of parallel threads

5
output_path str

Directory to save results

None
partial_frequency int

Save partial results every N entities

50
fail_safe bool

Retry previously failed entities

False
filter_column str

Column to filter by

None
filter_value any

Value to filter

None
filter_type str

'exclude' or 'include'

'exclude'
skip_export bool

Skip final export

False
prefix str

Output filename prefix

'vegetation_specific_dates'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Results DataFrame, errors, and statistics. When cache is active, also includes cache_hit and cache_miss counts.

MRTSExtractor 🔗

Bases: BaseExtractor

Extracts Medium Resolution Time Series (MRTS) vegetation data for agricultural entities.

Retrieves multi-sensor vegetation index time series using geometry-based queries. Supports raw and smoothed data extraction, temporal consistency checks, denoising, and end-of-curve extrapolation. Works with Sentinel-2, Landsat-8/9, and other sensors.

Documentation: https://docs.earthdaily.com/agro/library/Vegetation_time_series/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_MRTS_extraction_functions.ipynb

Args (setup_mrts_parameters): start_date (str): Start date in YYYY-MM-DD format. Default: '2025-05-01' end_date (str): End date in YYYY-MM-DD format. Default: '2025-10-15' sensors (list): Sensor list (Sentinel_2, Landsat_8, Landsat_9, etc.). Default: None (all) vegetation_index (str): Index type ('NDVI', 'EVI', 'CVI', 'GNDVI', 'NDWI', 'LAI'). Default: 'NDVI' aggregation (str): Aggregation method ('average', 'median', 'max', 'min'). Default: 'average' smoothing_method (str): Smoothing ('Whittaker', 'SavitzkyGolay', 'None'). Default: 'Whittaker' apply_denoiser (bool): Apply denoising filter. Default: True apply_end_of_curve (bool): Extrapolate end of curve. Default: True clear_cover_min (int): Minimum clear sky percentage (0-100). Default: 100 mask (str): Cloud mask used to discard NotClear pixels ('Native', 'ACM', 'Auto', 'ML', 'MLCirrus'). 'All' is NOT accepted here (coverage-only). Default: 'Auto' — best mask picked per image. Pass None to omit the key entirely (same behaviour). mode (str): Output mode ('full' = raw+smoothed, 'raw' = raw only). Default: 'full' compute_temporal_consistency (bool): Compute consistency checks. Default: True temporal_consistency_threshold (dict): Per-index thresholds. Example: {"Ndvi": 0.06, "Lai": 0.3} output_saturation (bool): Include index saturation flags in the output. Default: True extract_raw_datasets (bool): Include the raw per-image datasets alongside smoothed values. Default: True historical_years (int): Number of prior years to fetch for historical comparison. Default: 10 kpi_filter (dict): KPI aggregation rules applied to the time series. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop (required for LAI); start_date, end_date (optional overrides)

Output columns

entity_id, date, raw_value, noised, temporalConsistencyCheck, image_id, smoothed_value (full mode)

__init__ 🔗

__init__(bearer_token, token_expiration, config, workflow_ref=None)

Initialize Vegetation Time Series Extractor

Parameters:

Name Type Description Default
bearer_token str

Bearer token for API authentication

required
token_expiration datetime

Token expiration datetime

required
config dict

Configuration dictionary with env, paths, etc.

required
workflow_ref optional

Reference to WorkflowManager for token refresh

None

get_new_token 🔗

get_new_token()

Implements token refresh logic for cropidExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_mrts_parameters 🔗

setup_mrts_parameters(start_date='2025-05-01', end_date='2025-10-15', sensors=None, vegetation_index='NDVI', aggregation='average', smoothing_method='Whittaker', apply_denoiser=True, apply_end_of_curve=True, clear_cover_min=100, mask='Auto', output_saturation=True, extract_raw_datasets=True, compute_temporal_consistency=True, temporal_consistency_threshold=None, mode='full', historical_years=10, partial_frequency=50, kpi_filter=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for MRTS (Multi-Resolution Time Series) extraction.

Parameters:

Name Type Description Default
start_date str

Start date of the period (YYYY-MM-DD). Default: '2025-05-01'

'2025-05-01'
end_date str

End date of the period (YYYY-MM-DD). Default: '2025-10-15'

'2025-10-15'
sensors list

List of sensors to use. Available: Sentinel_2, Landsat_8, Landsat_9, HJ2B_CCD4, GAOFEN_6_WFV3 Default: None (omitted from payload → API uses all available sensors)

None
vegetation_index str

Vegetation index ('NDVI', 'EVI', 'CVI', 'GNDVI', 'NDWI', 'LAI'). Default: 'NDVI'

'NDVI'
aggregation str

Aggregation method ('average', 'median', 'max', 'min'). Default: 'average'

'average'
smoothing_method str

Smoothing method ('Whittaker', 'SavitzkyGolay', 'None'). Default: 'Whittaker'

'Whittaker'
apply_denoiser bool

Apply denoising. Default: True

True
apply_end_of_curve bool

Apply end of curve extrapolation. Default: True

True
clear_cover_min int

Minimum clear sky coverage percentage (0-100). Default: 100

100
mask str

Cloud mask applied before aggregation (TimeSeriesMaskType). One of 'Native', 'ACM', 'Auto', 'ML', 'MLCirrus' (case-insensitive; normalized to the API casing). Unlike the coverage/catalog-imagery endpoint, 'All' is rejected by this endpoint (HTTP 400) — it is a coverage-only value. Default: 'Auto' — the API selects the best available mask per image, which gives the most complete series. This is also what the API does when the key is absent, so the default is explicit, not a behaviour change; pass mask=None to omit it. Notes from live probing (prod, 2026-08-25, 3 sites x 2 seasons): - omitting mask is identical to mask='Auto' (same observations, not just the same count) — it is NOT Native and NOT ML - 'Auto' picks the best mask per image, so a single response can mix Native/ACM/ML labels; the 'mask' field on each raw point reports the mask actually applied, not the one requested - on fields where Auto happens to resolve to one mask throughout, omitted / 'Auto' / that mask are indistinguishable — compare observations, not counts - 'ACM' returns noticeably fewer raw observations than Native/ML - 'MLCirrus' is accepted but returned no data on any field tested

'Auto'
output_saturation bool

Include saturation information. Default: True

True
extract_raw_datasets bool

Extract raw datasets in addition to smoothed data. Default: True

True
compute_temporal_consistency bool

Compute temporal consistency metrics. Default: True

True
temporal_consistency_threshold dict

Threshold values for temporal consistency by index. Example: {"Ndvi": 0.06, "Lai": 0.3, "S2Rep": 2.5} If None and compute_temporal_consistency=True, uses default thresholds.

None
mode str

Output mode. Default: 'full' - 'full': Returns both raw and smoothed data merged on date - 'raw': Returns only raw observation data (no smoothed interpolation)

'full'
partial_frequency int

Partial save frequency during bulk processing. Default: 50

50
kpi_filter dict

KPI configuration for aggregating time series data: - 'kpi_name' (str): Name of the KPI - 'aggregation' (str): Type of aggregation (accumulation, average, max, min, count_gt, count_lt, count_between, std) - 'threshold' (float or tuple, optional): Threshold value(s) for count operations - Single float for count_gt/count_lt - Tuple of two floats (min, max) for count_between

None

get_mrts_api 🔗

get_mrts_api(entity_data: dict)

Get MRTS (Medium Resolution Time Series) data from API using geometry.

Parameters:

Name Type Description Default
entity_data dict

Entity data containing: - 'geometry' (str): WKT geometry string - 'id' (str, optional): Season field ID - 'crop' (str, optional): Crop type - 'sowing_date' (str, optional): Sowing date in YYYY-MM-DD format - 'start_date' (str, optional): Start date in YYYY-MM-DD format - 'end_date' (str, optional): End date in YYYY-MM-DD format If start_date/end_date not provided, uses dates from mrts_params

required

Returns:

Name Type Description
dict

JSON response with MRTS data

get_mrts_api_safe 🔗

get_mrts_api_safe(entity_data: dict) -> dict

Safe wrapper around get_mrts_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain 'geometry' and optionally 'id', 'start_date', 'end_date'

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str

dict

}

format_mrts_json 🔗

format_mrts_json(response_mrts_json, entity_data: dict = None)

Process MRTS API response into a DataFrame with date, raw and smoothed values.

Parameters:

Name Type Description Default
response_mrts_json dict

JSON response with 'rawData' and 'smoothedData' keys

required
entity_data dict

Original entity data to include ID in output

None

Returns:

Type Description

pd.DataFrame: - mode='full': entity_id (optional), date, raw_value, noised, temporalConsistencyCheck, mask, coveragePercent, image_id, smoothed_value - mode='raw': entity_id (optional), date, raw_value, noised, temporalConsistencyCheck, mask, coveragePercent, image_id

process_single_entity_mrts 🔗

process_single_entity_mrts(row, params=None)

Process a single entity for medium resolution time series extraction with optional KPI filtering.

Parameters:

Name Type Description Default
row pd.DataFrame, pd.Series, or dict

Entity data with required fields: - 'id': Entity identifier - 'geometry': WKT geometry string - 'start_date' (optional): Start date (YYYY-MM-DD) - falls back to params - 'end_date' (optional): End date (YYYY-MM-DD) - falls back to params - 'years' (optional): List of historical years for KPI comparison If DataFrame with multiple rows: only first row is processed

required
params dict

Override mrts_params (includes kpi_filter config and fallback dates)

None

Returns:

Name Type Description
dict

Contains 'data' (DataFrame) and 'error' - If KPI requested: DataFrame with metadata + KPI values (one row per entity) - If NO KPI: DataFrame with metadata + daily date/value data (multiple rows)

process_mrts_bulk_parallel 🔗

process_mrts_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='mrts', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity medium resolution vegetation time series using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain 'id' and 'geometry'.

required
params dict

Vegetation time series parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "mrts".

'mrts'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.

WeatherExtractor 🔗

Bases: BaseExtractor

Extracts field-level weather data for agricultural entities.

Retrieves historical daily weather data and agro-meteorological indices for a given geometry and date range. Supports configurable weather parameters and KPI aggregation.

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_weather.ipynb

Args (setup_weather_parameters): weather_type (str): Weather data type ('HISTORICAL_DAILY'). Default: 'HISTORICAL_DAILY' weather_parameters (str|list): Parameters to retrieve; 'none' expands to all available parameters. Default: 'precipitation.cumulative' kpi_filter (dict): KPI aggregation config (kpi_name, aggregation, threshold) historical_years (int): Number of prior years of weather history to fetch. Default: 0 page_limit (int): Override the API row limit ($limit); by default it is auto-sized to the query span so multi-year ranges are not truncated. Default: None use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); start_date, end_date (optional overrides)

Output columns

entity_id, date, + dynamic weather parameter columns (temperature, precipitation, etc.)

get_new_token 🔗

get_new_token()

Implements token refresh logic for WeatherExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_weather_parameters 🔗

setup_weather_parameters(weather_type='HISTORICAL_DAILY', weather_parameters='precipitation.cumulative', historical_years=0, partial_frequency=50, kpi_filter=None, page_limit=None, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure params for weather extraction for an entity

Parameters:

Name Type Description Default
weather_type str

Type of weather data (HISTORICAL_DAILY, FORECAST_DAILY, FORECAST_HOURLY)

'HISTORICAL_DAILY'
weather_parameters str or list

Weather parameters to extract. Defaults to 'precipitation.cumulative'. Pass 'none' to retrieve every available parameter.

'precipitation.cumulative'
historical_years int or list

Number of historical years for KPI comparison, or list of specific years

0
partial_frequency int

How often to save partial results

50
kpi_filter dict

KPI configuration

None
page_limit int

Override the API row limit ($limit). By default the limit is auto-sized to the query span (see get_weather_data) so long multi-year ranges are not truncated; set this to force an explicit value. Must be a positive integer. Default: None (auto-size).

None

get_weather_data 🔗

get_weather_data(entity_data)

Get weather data for selected parameters using self.weather_params

The API returns one row per day in ascending date order and caps the response at $limit rows, so a fixed limit silently truncates long ranges to the earliest N days (dropping the most recent data — the current-period window for KPIs). The row limit is therefore auto-sized to the query span via api_utils.autosize_row_limit (floor 1000 keeps short single-season/forecast requests byte-identical). A page_limit set in setup_weather_parameters wins over the auto-sized value.

Parameters:

Name Type Description Default
entity_data dict

Entity data with geometry, start_date, end_date

required

Returns:

Name Type Description
dict

Weather API response

get_weather_data_safe 🔗

get_weather_data_safe(entity_data: dict) -> dict

Safe wrapper around get_weather_data(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain 'id', 'geometry', 'start_date', 'end_date' keys

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str

dict

}

format_weather_json 🔗

format_weather_json(response_weather_json, entity_data: dict = None)

Format weather API response into a DataFrame. Dynamically handles all weather parameters returned by the API.

Parameters:

Name Type Description Default
response_weather_json list

JSON response from get_weather_data

required
entity_data dict

Original entity data to include ID in output

None

Returns:

Type Description

pd.DataFrame: DataFrame with date and all weather parameter columns

process_single_entity_weather 🔗

process_single_entity_weather(row, params=None)

Process a single entity for weather data extraction with optional KPI filtering.

Parameters:

Name Type Description Default
row pd.DataFrame, pd.Series, or dict

Entity data with required fields

required
params dict

Override weather_params

None

Returns:

Name Type Description
dict

Contains 'data' (DataFrame) and 'error'

process_entity_weather_bulk_parallel 🔗

process_entity_weather_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='weather', generate_report=False, report_options=None, use_cache=None, spatial_grouping=False, spatial_precision=DEFAULT_SPATIAL_PRECISION, spatial_max_window_days=DEFAULT_MAX_WINDOW_DAYS)

Bulk processing of entity weather data using threads + progress bar, with optional fail-safe retry and filter capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process

required
params dict

Weather parameters override

None
max_workers int

Number of threads to use

5
output_path str

Directory to save results

None
partial_frequency int

How often to save partial results

50
fail_safe bool

If True, retries only previously failed entities

False
filter_column str

Column name to filter entities by

None
filter_value any

Value to filter

None
filter_type str

'exclude' or 'include'

'exclude'
merge_existing str

Merge strategy

None
skip_export bool

If True, skip final export

False
prefix str

Prefix for output filenames

'weather'
use_cache bool

Enable/disable caching for this call

None
spatial_grouping bool

If True, group fields by geohash(centroid) × historical_years, call the API once per group over the union of the group's date windows, and slice each member out of that single response. Weather is queried by field centroid (Location=<centroid>), so neighbouring fields return identical values and dedup safely; members are cut back to their own window, so the output matches an ungrouped run row for row. Weather parameters are all per-day — including precipitation.cumulative, which despite its name is measured to be window-independent — so no value correction is applied. With use_cache, the geohash-keyed spatial cache is consulted per cell before any request — a cell whose window is already on disk costs no API call, whatever field set asks for it.

False
spatial_precision int

Geohash precision for the bucket (default DEFAULT_SPATIAL_PRECISION).

DEFAULT_SPATIAL_PRECISION
spatial_max_window_days int

Cap on a group's widened window (default DEFAULT_MAX_WINDOW_DAYS). Cells whose union exceeds it are split into sub-groups, so one long-history field cannot force a decade-long pull on every field sharing its cell.

DEFAULT_MAX_WINDOW_DAYS

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics

cropidExtractor 🔗

Bases: BaseExtractor

Extracts crop identification analytics for agricultural entities.

AI-powered crop type classification using multi-temporal satellite imagery. Supports end-of-season and in-season mask types, historical crop rotation analysis, and configurable confidence thresholds.

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_cropid.ipynb

Args (setup_cropid_parameters): begin_year (int): Start year for historical analysis. Default: 2020 end_year (int): End year for analysis. Default: 2025 mask_type (str): Mask type ('EndSeason', 'InSeason', 'PreSeason'). Default: 'EndSeason' limit_nb_crop (int): Max number of crop candidates per field. Default: 1 crop_mask_percent (int): Minimum crop mask percentage. Default: 50 mode (str): Output mode — 'history' (one row per (year, crop)), 'year' (filter to entity crop, long-form), 'historical_season' (comma-separated string of matching years), 'full_history' (wide-form, one row per entity, one column per year in begin_year..end_year filled with the top crop code per year). Default: 'history' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required); crop (required for 'year' and 'historical_season' modes)

Output columns

Depend on mode: - 'history' : entity_id, year, eda_crop_code, crop_name, raw_crop_name, crop_mask_percent - 'year' : entity_id, year, eda_crop_code - 'historical_season' : entity_id, historical_season (comma-separated years) - 'full_history' : entity_id + one column per year in [begin_year, end_year], filled with the top crop code (highest cropMaskPercent) per year

get_new_token 🔗

get_new_token()

Implements token refresh logic for cropidExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_cropid_parameters 🔗

setup_cropid_parameters(begin_year=2020, end_year=2025, mask_type='EndSeason', limit_nb_crop=1, crop_mask_percent=50, mode='history', partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for Crop ID extraction.

Parameters:

Name Type Description Default
begin_year int

Starting year (YYYY format, e.g., 2020)

2020
end_year int

Ending year (YYYY format, e.g., 2025)

2025
mask_type str

Type of crop mask. Options: 'InSeason', 'EndSeason', 'PreSeason'

'EndSeason'
limit_nb_crop int

Maximum number of crops to return per field

1
crop_mask_percent int

Minimum percentage of field covered by crop (0-100)

50
mode str

Processing mode — one of: - 'history' : all crops/years (long-form, one row per (year, crop)) - 'year' : filtered to entity crop (long-form) - 'historical_season' : comma-separated string of matching years - 'full_history' : wide-form pivot, one row per entity with one column per year (begin_year..end_year) filled with the top crop code per year

'history'
partial_frequency int

How often to save partial results during bulk processing

50

Returns:

Type Description

None

get_cropid_api 🔗

get_cropid_api(entity_data: dict)

Request crop id for an entity

Parameters:

Name Type Description Default
entity_data dict

entity data (id, geometry, crop)

required

Returns:

Name Type Description
dict

JSON response from the crop ID API

get_cropid_api_safe 🔗

get_cropid_api_safe(entity_data: dict) -> dict

Safe wrapper around get_cropid_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "entity_id": str

dict

}

format_cropid_json 🔗

format_cropid_json(response_json, entity_data: dict = None)

Format crop ID API response into a DataFrame.

Parameters:

Name Type Description Default
response_json dict

JSON response from get_cropid_api

required
entity_data dict

Original entity data to include ID in output

None

Returns:

Type Description

pd.DataFrame: DataFrame with columns: entity_id (optional), year, crop_name, raw_crop_name, crop_mask_percent, eda_crop_code

format_cropid_year_json 🔗

format_cropid_year_json(response_json, entity_data: dict)

Format crop ID API response and filter to only years matching the entity's crop.

Parameters:

Name Type Description Default
response_json dict

JSON response from get_cropid_api

required
entity_data dict

Original entity data with 'crop' as string (e.g., {'crop': 'SOYBEANS'})

required

Returns:

Type Description

pd.DataFrame: DataFrame with entity_id, year, eda_crop_code for matching crops only

format_cropid_historical_season_json 🔗

format_cropid_historical_season_json(response_json, entity_data: dict) -> str

Format crop ID API response to return a comma-separated string of years where the crop matches the entity's crop.

Parameters:

Name Type Description Default
response_json dict

JSON response from get_cropid_api

required
entity_data dict

Entity data with 'crop' as string (e.g., {'crop': 'SOYBEANS'})

required

Returns:

Name Type Description
str str

Comma-separated years matching the entity crop (e.g., "2020,2022,2024,2025")

format_cropid_full_history_json 🔗

format_cropid_full_history_json(response_json, entity_data: dict)

Format crop ID API response into a wide-form, single-row DataFrame: one column per year in [begin_year, end_year], filled with the top crop code (highest cropMaskPercent) for that year — or NaN if the year is absent from the response.

Year columns are string-typed (e.g. "2020") so the wide frame plays nicely with parquet and bulk concat across entities. When multiple crops are reported for one year (limit_nb_crop > 1 upstream), only the top crop by cropMaskPercent is kept; tie-break is the API's natural ordering.

Parameters:

Name Type Description Default
response_json dict

JSON response from get_cropid_api

required
entity_data dict

Entity data — must include 'id'

required

Returns:

Type Description

pd.DataFrame: 1-row DataFrame with columns entity_id + one column per year between begin_year and end_year inclusive.

process_single_entity_cropid 🔗

process_single_entity_cropid(row, params=None)

Process a single entity for crop ID extraction.

Parameters:

Name Type Description Default
row dict

Entity data with 'id', 'geometry', and 'crop'

required
params dict

Override cropid_params

None

Returns:

Name Type Description
dict

Contains 'data' (DataFrame) and 'error' (dict or None)

process_cropid_bulk_extraction_parallel 🔗

process_cropid_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='cropid', generate_report=False, report_options=None, use_cache=None)

Bulk processing of cropid requests using threads + progress bar, with optional fail-safe retry, partial save capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id', 'geometry', and 'crop').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "cropid".

'cropid'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

{ "results_df": pd.DataFrame, "global_errors": list[dict], "total_entities": int, "total_calculations": int, "successful_calculations": int, "failed_calculations": int, "failed_ids": list[str] }

LocationBasedBorderExtractor 🔗

Bases: BaseExtractor

Extracts a field-border polygon from a single longitude/latitude location.

Calls the Geosys /field-borders/v1/AutomaticBoundary endpoint with a location=lon,lat query parameter and returns the polygon (in WKT) of the field that contains that point. Useful for converting a list of geocoded locations (CSV / API output) into proper field geometries.

Documentation: https://api.geosys-na.net/field-borders/v1/swagger/index.html Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_LocationBasedBorder.ipynb

Args (setup_location_based_border_parameters): simplified_geom (bool): If True, returns a simplified field geometry (lower shape-point count). Default: True. Maps to the API's simplified_geom query parameter. partial_frequency (int): How often to flush partial results. Default: 50. use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id (required) - used to stamp results, log lines, and cache keys. geometry (required) - must be a Point WKT (e.g. POINT(-93.6 41.5)). If the row carries a non-point geometry the entity is rejected with a clean error rather than being silently degraded to a centroid.

Output columns

entity_id, point_geometry (echo of input point as WKT), polygon_geometry (returned field border as WKT), plus any flat scalar fields the API returns alongside geometry (e.g. area, sourceId).

get_new_token 🔗

get_new_token()

Implements token refresh logic for LocationBasedBorderExtractor.

setup_location_based_border_parameters 🔗

setup_location_based_border_parameters(simplified_geom: bool = True, partial_frequency: int = 50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for location-based-border extraction.

Parameters:

Name Type Description Default
simplified_geom bool

If True, returns simplified field geometry. Default: True.

True
partial_frequency int

How often to flush partial results. Default: 50.

50
column_mapping dict

Column name overrides (e.g. {"geometry": "point_wkt"}).

None
output_mapping dict

Output column renames.

None
exclude_columns list

Columns to exclude from output.

None
output_columns list

Whitelist of output columns.

None
use_cache bool

Whether to use caching for this run.

None

get_location_based_border_api 🔗

get_location_based_border_api(entity_data)

Call the AutomaticBoundary API for a single entity.

Parameters:

Name Type Description Default
entity_data dict | Series

Entity data with id and geometry (a POINT WKT, e.g. POINT(-93.6 41.5)).

required

Returns:

Name Type Description
dict

Parsed JSON response from the API. Always carries the input

id and point_wkt echoed back via setdefault so the

downstream formatter can stamp them.

get_location_based_border_api_safe 🔗

get_location_based_border_api_safe(entity_data)

Safe wrapper around get_location_based_border_api().

Returns:

Name Type Description
dict

{"success": bool, "data": dict|None, "error": str|None, "entity_id": str}

format_location_based_border_json 🔗

format_location_based_border_json(response_json, params=None)

Normalize an AutomaticBoundary response into a single-row DataFrame.

Accepts three shapes (the API's response schema isn't documented in the Swagger so we tolerate the common variants):

  1. Flat dict with a geometry field — string WKT or GeoJSON geometry.
  2. GeoJSON Feature: {"type": "Feature", "geometry": {...}, "properties": {...}}.
  3. GeoJSON FeatureCollection — uses features[0].

Parameters:

Name Type Description Default
response_json dict

API response (augmented with id and point_wkt by get_location_based_border_api).

required
params dict

Override params (unused here, kept for symmetry).

None

Returns:

Type Description

pd.DataFrame: Single-row DataFrame with entity_id,

point_geometry, polygon_geometry, and any flat scalar

properties returned by the API.

process_single_entity_location_based_border 🔗

process_single_entity_location_based_border(row, params=None)

Process location-based-border extraction for a single entity with retry logic.

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_location_based_border_bulk_extraction_parallel 🔗

process_location_based_border_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='location_based_border', generate_report=False, report_options=None, use_cache=None)

Bulk processing of location-based-border requests using threads + progress bar.

Args mirror the family pattern (coverage_function / difference_functions / processor_change_index_functions). Returns the standard {results_df, global_errors, total_entities, total_calculations, successful_calculations, failed_calculations, failed_ids} dict.

ZoningExtractor 🔗

Bases: BaseExtractor

Extracts management zone maps (SAMZ) for agricultural entities.

Calls the SAMZ (SAtellite derived Management Zones) API to get management zones from satellite imagery. Returns field-level statistics including variability, productivity indices, and per-zone area/productivity breakdowns.

Documentation: https://docs.earthdaily.com/agro/library/Field%20Level%20Maps/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_zoning.ipynb

Args (setup_zoning_parameters): num_zones (int): Number of management zones to generate. Default: 5 output_epsg (int): Output coordinate system. Default: 4326 postprocess (str): Processing mode - 'stats', 'stats_geo', 'links', or 'file'. Default: 'stats' 'stats' returns one field-level row; 'stats_geo' returns one row per zone, carrying that zone's geometry alongside the field-level stats. map_format (str): Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None output_path (str): Directory for file downloads. Required when postprocess='file'. directLinks (bool): Request direct download links from API. Default: False partial_frequency (int): How often to save partial results. Default: 50 use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry, image_id (required)

Output columns

Varies by postprocess mode. stats: entity_id, field_variability, most_variable_zone, highest_interzone_variability, field_productivity_index, field_variability_index, zone_{n}area_percent, zoneproductivity_index, zonevariability_index stats_geo: one row per zone — zone_name, zone_geometry, productivity_index, variability_index, area_percent, number_of_pixels, zone_mean, zone_max, zone_min, zone_area, plus the field-level stats above repeated on every row links: image_png_link, worldfile_link, thumbnail_link, worldfile, bbox_, map_width, map_height file: status, map_format, file_count, total_size_bytes, saved_files

TODO
  • Cache support: cache is disabled by default because results depend on image_id (which can be a list). When image_id is a list, it gets stored as a list object in the DataFrame, which does not deduplicate cleanly in parquet. To enable cache properly, image_id lists need to be serialized to a stable string (e.g. sorted, joined with ";") before cache lookup and storage, so that the same set of images always produces the same cache key regardless of order.

get_new_token 🔗

get_new_token()

Implements token refresh logic for ZoningExtractor.

setup_zoning_parameters 🔗

setup_zoning_parameters(num_zones=5, output_epsg=4326, postprocess='stats', map_format=None, output_path=None, skip_existing=True, directLinks=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for management zone extraction.

Parameters:

Name Type Description Default
num_zones int

Number of management zones (2-10). Default: 5

5
output_epsg int

Output EPSG code for projection. Default: 4326

4326
postprocess str

Processing mode - 'stats', 'stats_geo', 'links', or 'file'. Default: 'stats' 'stats' returns a single field-level row per entity. 'stats_geo' returns one row per zone, adding the zone geometry (segments merged into a GEOMETRYCOLLECTION when a zone has several) and its per-zone mean/max/min/area.

'stats'
map_format str

Output format for file mode ('png', 'tiff.zip', 'shp.zip'). Default: None

None
output_path str

Directory for file downloads. Required when postprocess='file'.

None
skip_existing bool

Skip download if file already exists. Default: True

True
directLinks bool

Request direct download links. Default: False Note: Automatically set to True when postprocess='links'

False
partial_frequency int

How often to save partial results. Default: 50

50
column_mapping dict

Column name overrides.

None
output_mapping dict

Output column renames.

None
exclude_columns list

Columns to exclude from output.

None
output_columns list

Whitelist of output columns.

None
use_cache bool

Whether to use caching.

None

get_zoning_map_api 🔗

get_zoning_map_api(entity_data: dict)

Call the SAMZ API to generate management zones for an entity. Returns raw Response when map_format is set (file download), or parsed JSON otherwise (stats/links).

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id', 'geometry', and 'image_id' image_id can be a single string or a list of image ID strings.

required

Returns:

Type Description

requests.Response (file mode) or dict (stats/links mode)

get_zoning_map_api_safe 🔗

get_zoning_map_api_safe(entity_data: dict) -> dict

Safe wrapper around get_zoning_map_api().

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id', 'geometry', and 'image_id'

required

Returns:

Name Type Description
dict dict

{"success": bool, "data": dict|None, "error": str|None, "seasonfield_id": str}

format_zoning_stats_json 🔗

format_zoning_stats_json(response_json)

Extract field-level and per-zone statistics from SAMZ API response.

Parameters:

Name Type Description Default
response_json dict

Parsed API response

required

Returns:

Type Description

pd.DataFrame: Single-row DataFrame with field stats and per-zone columns

format_zoning_stats_geo_json 🔗

format_zoning_stats_geo_json(response_json)

Extract per-zone statistics with geometry, productivity index, and variability index.

Produces one row per zone, combining data from legend.ranges (productivity/variability) and zones (geometry/stats). Zone geometries from multiple segments are merged.

Parameters:

Name Type Description Default
response_json dict

Parsed API response

required

Returns:

Type Description

pd.DataFrame: One row per zone with geometry, indices, and field-level stats

format_zoning_links_json(response_json)

Extract direct links and metadata from SAMZ API response.

Parameters:

Name Type Description Default
response_json dict or Response

API response with _links, worldFile, etc.

required

Returns:

Type Description

pd.DataFrame: Single-row DataFrame with links and spatial metadata

format_zoning_file 🔗

format_zoning_file(entity_data: dict, image_id, response, output_path=None)

Save SAMZ zone map file response (PNG, TIFF.ZIP, or SHP.ZIP) to disk.

Parameters:

Name Type Description Default
entity_data dict

Entity data including 'id'

required
image_id str or list

Image ID(s) for filename. If a list, joined with '_' for the filename.

required
response Response

Raw response from get_zoning_map

required
output_path str

Directory to save files.

None

Returns:

Name Type Description
dict

Status and file path information

process_single_entity_zoning 🔗

process_single_entity_zoning(row, params=None)

Process zoning extraction for a single entity with retry logic. Routes to stats, links, or file processing based on postprocess parameter.

Parameters:

Name Type Description Default
row dict | Series

Entity data with id, geometry, image_id

required
params dict

Override parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_entity_zoning_bulk_parallel 🔗

process_entity_zoning_bulk_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='zoning', generate_report=False, report_options=None, use_cache=None)

Bulk processing of entity SAMZ zones using threads + progress bar, with optional fail-safe retry and filter capabilities.

Parameters:

Name Type Description Default
entity_list DataFrame

List of entities to process, must contain 'id', 'geometry', and 'image_id'.

required
params dict

Zoning parameters override.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final CSV results.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export. Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "zoning".

'zoning'
generate_report bool

Generate HTML report. Default: False.

False
report_options dict

Report configuration options.

None
use_cache bool

Enable/disable caching for this call.

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics.