Skip to content

Sustainability extractors🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

BaresoilExtractor 🔗

Bases: BaseExtractor

Extracts bare soil exposure analytics for agricultural entities (experimental).

Estimates the number of days with bare soil exposed during a season using satellite imagery analysis. Supports summary and detailed output modes.

Documentation: https://docs.earthdaily.com/agro/library/baresoil/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_baresoil.ipynb

Args (setup_baresoil_parameters): season_duration (int): Season length in days. Default: 120 season_start_day (int): Season start day of month. Default: 1 season_start_month (int): Season start month. Default: 4 year (int): Target year. Default: 2025 filter (str): Output mode ('summary' or 'full'). Default: 'summary' use_cache (bool): Reuse cached API responses and cache new results to avoid re-fetching. None uses the extractor's instance default. Default: None

Entity fields (via column_mapping): id, geometry (required)

Output columns

entity_id, bare_soil_days, + baresoil detection metrics

get_new_token 🔗

get_new_token()

Implements token refresh logic for BaresoilExtractor. Called automatically by BaseExtractor.ensure_token_valid() if token is expired.

setup_baresoil_parameters 🔗

setup_baresoil_parameters(season_duration=120, season_start_day=1, season_start_month=4, year=2025, filter='summary', publish_af=False, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for baresoil extraction

Parameters:

Name Type Description Default
season_duration int

Crop cycle duration in days

120
season_start_day int

Season start day (1-31)

1
season_start_month int

Season start month (1-12)

4
year int

Year to analyze

2025
publish_af bool

Whether to publish to AF (includes id in payload if True). Default: False

False
partial_frequency int

How often to save partial results

50

get_baresoil_api 🔗

get_baresoil_api(entity_data: dict)

Request baresoil detection for an entity

Parameters:

Name Type Description Default
entity_data dict

entity data (id, geometry, year)

required

get_baresoil_api_safe 🔗

get_baresoil_api_safe(entity_data: dict) -> dict

Safe wrapper around get_baresoil_api(). Returns structured response with success flag, data or error.

Parameters:

Name Type Description Default
entity_data dict

Must contain a 'geometry' key in WKT format.

required

Returns:

Name Type Description
dict dict

{ "success": bool, "data": dict | None, "error": str | None, "seasonfield_id": str

dict

}

format_baresoil_json 🔗

format_baresoil_json(response_baresoil_json, params=None)

Normalize Baresoil API response into a pandas DataFrame.

Expected structure

{ 'id': '', 'data': { 'year': 2025, 'seasonDuration': 200, 'seasonStartDay': 1, 'seasonStartMonth': 1, 'baresoilDays': 48, 'baresoilPeriods': [ {'start': '2025-01-28', 'end': '2025-03-16', 'periodLength': 48} ] } }

Parameters:

Name Type Description Default
response_baresoil_json dict

API response JSON.

required
params dict

Baresoil parameters (defaults to self.baresoil_params).

None

Returns:

Type Description

pd.DataFrame: - If filter='summary': One row per entity with data values only - If filter='full': One row per baresoil period with all data and period details

process_single_entity_baresoil 🔗

process_single_entity_baresoil(row, params=None)

Process baresoil extraction for a single entity with retry logic.

Parameters:

Name Type Description Default
row dict

Entity data containing id and geometry

required
params dict

Baresoil parameters

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}

process_baresoil_bulk_extraction_parallel 🔗

process_baresoil_bulk_extraction_parallel(entity_list, params=None, max_workers=5, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='baresoil', generate_report=False, report_options=None, use_cache=None)

Bulk processing of baresoil requests using threads + progress bar, with optional fail-safe retry, filter capabilities, and caching.

Parameters:

Name Type Description Default
entity_list DataFrame

Entities to process (must contain 'id' and 'geometry').

required
params dict

Override monitoring parameters.

None
max_workers int

Number of threads to use.

5
output_path str

Directory to save final results and error logs.

None
partial_frequency int

How often to save partial results.

50
fail_safe bool

If True, per-entity failures are captured to failed_ids instead of aborting the run. Error tolerance only -- it does NOT change which entities are processed. To reprocess just the IDs a previous run recorded, set retry_failed_only (see BaseExtractor).

False
filter_column str

Column name to filter entities by.

None
filter_value any

Value to filter in the filter_column.

None
filter_type str

'exclude' to skip rows with filter_value, 'include' to process only rows with filter_value. Defaults to 'exclude'.

'exclude'
merge_existing str

Merge strategy - 'auto', 'preserve', or 'mark'. If None, uses instance default.

None
skip_export bool

If True, skip final export (useful when chaining extractions). Default: False.

False
prefix str

Prefix for output filenames and failed IDs files. Default: "baresoil".

'baresoil'
use_cache bool

Override instance-level cache setting. Default: None (uses self.use_cache).

None

Returns:

Name Type Description
dict

Contains results DataFrame, errors list, and summary statistics. When cache is active, also includes cache_hit and cache_miss counts.