Skip to content

Entity and user management🔗

Generated from the source docstrings by mkdocstrings at build time. See the API reference overview for the other groups.

EntityManager 🔗

Bases: BaseExtractor

Manages Farm, Field, and Seasonfield entities via EarthDaily MDM API.

Provides full CRUD operations for the agricultural entity hierarchy (Farm > Field > Seasonfield). Supports bulk entity loading, batch processing, and integration with UserManager for Grower ID resolution.

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/ Notebook: https://github.com/earthdaily/Examples-and-showcases/blob/main/agriculture/EDAgriculture_entity_management.ipynb

Supported operations
  • Farm: create, get, update, delete
  • Field: create, get, update, delete
  • Seasonfield: create, get, update, delete
  • Batch loading: load_seasonfields(), load_seasonfields_batch()
Entity fields

id, geometry, crop, sowing_date, name, farm_name (varies by operation)

__init__ 🔗

__init__(bearer_token, token_expiration, config, workflow_ref=None, user_manager: UserManager = None)

Initialize EntityManager.

Parameters:

Name Type Description Default
bearer_token str

API bearer token

required
token_expiration datetime

Token expiration time

required
config dict

Configuration dictionary

required
workflow_ref

Optional workflow manager reference

None
user_manager UserManager

UserManager instance for Grower ID resolution

None

get_new_token 🔗

get_new_token()

Refresh token using EDAuthenticator.

clear_farm_cache 🔗

clear_farm_cache()

Clear the farm name -> ID cache.

load_input_file 🔗

load_input_file(file_path: str, encoding: str = 'utf-8') -> pd.DataFrame

Load input file with automatic format detection (CSV, Excel, Shapefile).

Expected columns: - Grower: Grower ID (alphanumeric) - Farm: Farm name - Seasonfield: Field/Seasonfield name - Sowing: Sowing date (YYYY-MM-DD) - Crop: Crop code (e.g., CORN) - Geometry: WKT geometry (for CSV/Excel) or auto-extracted (for SHP)

Parameters:

Name Type Description Default
file_path str

Path to input file

required
encoding str

File encoding for CSV (default: utf-8)

'utf-8'

Returns:

Type Description
DataFrame

pd.DataFrame: Loaded and normalized data

get_grower_id 🔗

get_grower_id(login: str) -> Optional[str]

Get Grower ID by login using UserManager.

Parameters:

Name Type Description Default
login str

User login (username, not email)

required

Returns:

Name Type Description
str Optional[str]

Grower ID (alphanumeric) or None if not found

get_growers 🔗

get_growers(fields: str = 'id,userType.code,companyName,firstname,lastname,login') -> pd.DataFrame

Get all GROWER users.

Parameters:

Name Type Description Default
fields str

Fields to retrieve

'id,userType.code,companyName,firstname,lastname,login'

Returns:

Type Description
DataFrame

pd.DataFrame: Growers data

get_crops 🔗

get_crops(fields: str = 'id,cycleduration') -> pd.DataFrame

Get all available crops.

Parameters:

Name Type Description Default
fields str

Fields to retrieve

'id,cycleduration'

Returns:

Type Description
DataFrame

pd.DataFrame: Crops data with id and properties

create_farm 🔗

create_farm(grower_id: str, farm_name: str, use_cache: bool = True) -> str

Create a Farm under a Grower. Uses cache to avoid duplicates.

Parameters:

Name Type Description Default
grower_id str

Grower user ID

required
farm_name str

Name of the farm

required
use_cache bool

Use internal cache to avoid duplicate creation

True

Returns:

Name Type Description
str str

Created or existing Farm ID

get_farms 🔗

get_farms(grower_id: str = None, fields: str = 'id,name,grower.id') -> pd.DataFrame

Get farms with optional grower filter.

get_farm_by_id 🔗

get_farm_by_id(farm_id: str) -> Optional[Dict]

Get single farm by ID.

update_farm 🔗

update_farm(farm_id: str, farm_name: str = None, grower_id: str = None) -> bool

Update a Farm.

delete_farm 🔗

delete_farm(farm_id: str) -> bool

Delete a Farm.

create_field 🔗

create_field(farm_id: str, field_name: str, geometry: str) -> str

Create a Field under a Farm.

get_fields 🔗

get_fields(farm_id: str = None, fields: str = 'id,name,geometry,farm.id') -> pd.DataFrame

Get fields with optional farm filter.

get_field_by_id 🔗

get_field_by_id(field_id: str) -> Optional[Dict]

Get single field by ID.

update_field 🔗

update_field(field_id: str, field_name: str = None, geometry: str = None, farm_id: str = None) -> bool

Update a Field.

delete_field 🔗

delete_field(field_id: str) -> bool

Delete a Field.

create_seasonfield 🔗

create_seasonfield(field_id: str, geometry: str, sowing_date: str, crop_code: str = 'CORN', name: str = None) -> str

Create a Seasonfield under a Field.

get_seasonfields 🔗

get_seasonfields(field_id: str = None, crop_code: str = None, sowing_date_gte: str = None, sowing_date_lte: str = None, farm_name: str | list[str] | None = None, external_ids: dict | None = None, fields: str = 'id,name,geometry,sowingDate,crop.code,field.id,crop,field.farm.grower.companyName,field.farm.grower.firstname') -> pd.DataFrame

Get seasonfields with optional filters.

Parameters:

Name Type Description Default
field_id str

Filter by field ID

None
crop_code str

Filter by crop code (e.g., 'CORN', 'SOYBEANS')

None
sowing_date_gte str

Filter sowingDate >= this date (YYYY-MM-DD)

None
sowing_date_lte str

Filter sowingDate <= this date (YYYY-MM-DD)

None
farm_name str | list[str]

Filter by farm name (Field.Farm.Name). Pass a single name ("My Farm") or a list (["Farm A", "Farm B"]) to filter several farms in ONE request via the MDM $in: operator.

None
external_ids dict

Filter by one or more externalIds systems. Maps an externalIds sub-key to a value or list of values, e.g. {"smbsC_ID": ["7073", "7074"]} -> externalIds.smbsC_ID=$in:7073|7074. Keys may be given bare ("smbsC_ID") or fully-qualified ("externalIds.smbsC_ID"). To get the externalIds back in the response, also include externalIds in fields.

None
fields str

Comma-separated list of fields to retrieve. Include externalIds to retrieve the nested external-id object (flattened to externalIds.* columns on the way out).

'id,name,geometry,sowingDate,crop.code,field.id,crop,field.farm.grower.companyName,field.farm.grower.firstname'

Returns:

Type Description
DataFrame

pd.DataFrame: DataFrame with seasonfield data (deduplicated on id). When externalIds is requested its nested object is flattened to externalIds.<system> columns; if the native id field was not requested it is recovered from externalIds.id so the frame stays usable by the extractors.

Example

mgr.get_seasonfields(farm_name=["Brazil Demo Farm", "France Demo Farm"]) mgr.get_seasonfields( external_ids={"smbsC_ID": ["7073", "7074"]}, fields="externalIds,geometry,sowingDate,customerExternalId", )

load_seasonfields 🔗

load_seasonfields(sowing_date_gte=None, sowing_date_lte=None, crop_id=None, start_date=None, end_date=None, farm_name=None, external_ids=None, fields=None) -> pd.DataFrame

Load seasonfields from EarthDaily platform with optional filters.

Parameters:

Name Type Description Default
sowing_date_gte str

Filter sowingDate >= this date (YYYY-MM-DD)

None
sowing_date_lte str

Filter sowingDate <= this date (YYYY-MM-DD)

None
crop_id str

Filter by crop ID (e.g., 'CORN', 'SOYBEANS', 'WHEAT')

None
start_date str

Legacy alias for sowing_date_gte

None
end_date str

Legacy alias for sowing_date_lte

None
farm_name str | list[str]

Filter by farm name (Field.Farm.Name). A single name or a list of names (filtered in one request via $in:).

None
external_ids dict

Filter by externalIds systems, e.g. {"smbsC_ID": ["7073", "7074"]}. See get_seasonfields() for details.

None
fields str

Comma-separated list of fields to retrieve. If None, uses get_seasonfields default. Include externalIds to get the external-id object back (flattened to externalIds.* columns).

None

Returns:

Type Description
DataFrame

pd.DataFrame: Loaded seasonfields

Example

entity_manager.load_seasonfields(sowing_date_gte="2025-01-01", crop_id="CORN", farm_name="My Farm") entity_manager.load_seasonfields(farm_name=["Farm A", "Farm B"]) # both farms, one request entity_manager.load_seasonfields( external_ids={"smbsC_ID": ["7073", "7074"]}, fields="externalIds,geometry,sowingDate,customerExternalId", )

load_seasonfields_batch 🔗

load_seasonfields_batch(start_date=None, end_date=None, crop_id=None, ids=None, batch_size=5000) -> pd.DataFrame

Load seasonfields using batch processing for improved performance.

Parameters:

Name Type Description Default
start_date str

Start date for sowingDate filter (format: YYYY-MM-DD)

None
end_date str

End date for sowingDate filter (format: YYYY-MM-DD)

None
crop_id str

Crop ID to filter by

None
ids List[str]

List of specific seasonfield IDs to load. If None, loads all available seasonfields.

None
batch_size int

Number of seasonfields to process per batch (default: 5000)

5000

Returns:

Type Description
DataFrame

pd.DataFrame: DataFrame containing the loaded seasonfields

get_seasonfield_by_id 🔗

get_seasonfield_by_id(seasonfield_id: str) -> Optional[Dict]

Get single seasonfield by ID.

update_seasonfield 🔗

update_seasonfield(seasonfield_id: str, geometry: str = None, sowing_date: str = None, crop_code: str = None, name: str = None) -> bool

Update a Seasonfield.

delete_seasonfield 🔗

delete_seasonfield(seasonfield_id: str) -> bool

Delete a Seasonfield.

prepare_field_dataframe 🔗

prepare_field_dataframe(df: DataFrame, *, grower_id: str, farm_name: str, column_mapping: Optional[Dict[str, str]] = None, default_crop: str = 'OTHERS', default_sowing_date: Optional[str] = None, source_epsg: Optional[Any] = None, min_area_ha: Optional[float] = 1.0, field_name_template: str = 'Field {n}', verbose: bool = True) -> tuple

Turn an arbitrary field file into rows create_entities_from_dataframe accepts.

Geometry is the only thing the caller must supply. Everything else the entity hierarchy needs is either defaulted here or comes from the provisioning run, because in practice a client sends a shapefile or a CSV of borders and nothing else:

============= ========================================================== Grower always grower_id — the GROWER user this run created. Never read from the file: fields belong to the account being provisioned, and a stale id in a spreadsheet would silently file them under someone else. Farm always farm_name — the farm this run creates. Geometry required. Normalised by :func:~earthdaily.agriculture.core.geometry.normalize_field_geometry — reprojected to EPSG:4326, checked areal and valid, size-floored. Crop from the file when present, else default_crop. An unknown code also falls back, and is reported. Sowing from the file when present, else default_sowing_date. Seasonfield from the file when present, else generated from field_name_template. ============= ==========================================================

Nothing is dropped silently. Rows whose geometry cannot be salvaged come back in a second frame with the reason, so a 500-field import shows what it refused instead of quietly creating 480 fields.

Parameters:

Name Type Description Default
df DataFrame

Raw input, straight from :meth:load_input_file or any reader.

required
grower_id str

GROWER user id to file every field under.

required
farm_name str

Farm name to file every field under.

required
column_mapping Optional[Dict[str, str]]

Explicit {source column: canonical name}. Applied before the built-in aliases, so it always wins. Canonical names are Geometry / Seasonfield / Crop / Sowing.

None
default_crop str

Crop for rows without a usable one.

'OTHERS'
default_sowing_date Optional[str]

YYYY-MM-DD for rows without one. Required when the file has no sowing column.

None
source_epsg Optional[Any]

CRS of the input geometry. None means "assume 4326 and verify"; out-of-range coordinates then raise rather than guess.

None
min_area_ha Optional[float]

Size floor. None disables it.

1.0
field_name_template str

Used when no name column exists. {n} is the 1-based row number.

'Field {n}'
verbose bool

Print the summary.

True

Returns:

Name Type Description
tuple tuple

(ready, rejected). ready carries the six columns

tuple

meth:create_entities_from_dataframe requires plus Area_ha and

tuple

Geometry_notes; rejected carries the original row plus

tuple

Reason.

Raises:

Type Description
ValueError

No geometry column, or no sowing date available at all.

create_entities_from_dataframe 🔗

create_entities_from_dataframe(df: DataFrame, verbose: bool = True) -> pd.DataFrame

Create Farm -> Field -> Seasonfield hierarchy from DataFrame. Reuses existing Farm IDs when farm name repeats.

Parameters:

Name Type Description Default
df DataFrame

Input data with columns [Grower, Farm, Seasonfield, Sowing, Crop, Geometry]

required
verbose bool

Print progress

True

Returns:

Type Description
DataFrame

pd.DataFrame: Input data enriched with Farm_Id, Field_Id, Seasonfield_Id, Status, Details

delete_entities_from_dataframe 🔗

delete_entities_from_dataframe(df: DataFrame, delete_farms: bool = False, verbose: bool = True) -> pd.DataFrame

Delete Seasonfields and Fields from DataFrame. Optionally delete Farms (careful: affects all fields under farm).

Parameters:

Name Type Description Default
df DataFrame

Data with Seasonfield_Id, Field_Id, (optionally Farm_Id)

required
delete_farms bool

Also delete farms (default: False)

False
verbose bool

Print progress

True

Returns:

Type Description
DataFrame

pd.DataFrame: Input data enriched with Status, Details

export_results 🔗

export_results(results_df: DataFrame, operation: str, output_path: str = None) -> str

Export operation results to CSV.

Parameters:

Name Type Description Default
results_df DataFrame

Results DataFrame

required
operation str

Operation type (creation/modification/deletion/export)

required
output_path str

Output directory

None

Returns:

Name Type Description
str str

Path to exported file

UserManager 🔗

Bases: BaseExtractor

Manages user accounts via EarthDaily MDM API.

Handles user lifecycle operations including creation, role assignment, account linking, modification, and deletion. All batch operations use parallel processing for performance.

Documentation: https://docs.earthdaily.com/agro/library/Api_reference/

Supported operations
  • User creation (AGRONOMIST/GROWER) with automatic ordering
  • Role assignment (multiple roles)
  • Account linking (AGRONOMIST manages GROWER)
  • User modification (identified by ID) with link synchronization
  • User deletion with automatic link cleanup

__init__ 🔗

__init__(bearer_token, token_expiration, config, workflow_ref=None)

Initialize UserManager.

get_new_token 🔗

get_new_token()

Refresh token using EDAuthenticator.

load_csv 🔗

load_csv(file_path: str, operation: str = 'create', encoding: str = 'utf-8') -> pd.DataFrame

Load CSV with automatic separator detection.

get_users 🔗

get_users(user_type: Optional[str] = None, login: Optional[str] = None, customer_id: Optional[str] = None, fields: str = 'id,userType.code,companyName,firstname,lastname,email,login') -> pd.DataFrame

Get users from API with roles and relationships.

Parameters:

Name Type Description Default
user_type Optional[str]

Filter by user type code (e.g. "AGRONOMIST", "GROWER").

None
login Optional[str]

Filter by exact login.

None
customer_id Optional[str]

Filter by customer id — the usual back-office entry point for pulling one client's whole user base.

None
fields str

$fields projection sent to the API.

'id,userType.code,companyName,firstname,lastname,email,login'

Returns:

Type Description
DataFrame

pd.DataFrame with back-office column names (User_Type, First_Name,

DataFrame

Last_Name, CompanyName), a blank Password column positioned

DataFrame

after login, plus Role and Managed_by_user from the

DataFrame

relationship enrichment. Shape round-trips through load_csv().

get_user_by_login 🔗

get_user_by_login(login: str) -> Optional[Dict]

Get single user by login.

get_user_by_id 🔗

get_user_by_id(user_id: str) -> Optional[Dict]

Get single user by ID.

get_managed_growers 🔗

get_managed_growers(agronomist_id: str) -> List[Dict]

Get list of GROWER accounts managed by an AGRONOMIST.

setup_user_parameters 🔗

setup_user_parameters(fields=None, partial_frequency=50, column_mapping=None, output_mapping=None, exclude_columns=None, output_columns=None, use_cache=None)

Configure parameters for bulk user reads.

Parameters:

Name Type Description Default
fields list[str]

Subset of (flattened) user columns to keep in the output. None keeps every field returned by the API.

None
partial_frequency int

How often to flush partial bulk results.

50
column_mapping dict

Maps canonical entity fields to your DataFrame columns. The id mapping must point at the user id.

None
output_mapping / exclude_columns / output_columns

Output formatting.

required
use_cache bool

Per-call cache override.

None

format_user_json 🔗

format_user_json(response_json, entity_data=None)

Normalise a single user record (from get_user_by_id) into a one-row DataFrame, flattening nested objects (userType.code, country.code, ...).

Returns:

Type Description

pd.DataFrame | None: One-row frame, or None if the record is empty.

process_single_entity_user 🔗

process_single_entity_user(row, params=None)

Fetch one user by id with retry logic.

Parameters:

Name Type Description Default
row dict | Series | DataFrame

Entity data; the mapped id column holds the MDM user id.

required
params dict

Override for self.user_params.

None

Returns:

Name Type Description
dict

{"data": DataFrame or None, "error": dict or None}.

process_entity_user_bulk_parallel 🔗

process_entity_user_bulk_parallel(entity_list, params=None, max_workers=10, output_path=None, partial_frequency=50, fail_safe=False, filter_column=None, filter_value=None, filter_type='exclude', merge_existing=None, skip_export=False, prefix='users', generate_report=False, report_options=None, use_cache=None)

Bulk user read with threading, partial saves, cache and merge logic.

Conforms to the WorkflowManager run-method contract, so UserManager can be wired into a YAML workflow as an extractor step::

run: {method: process_entity_user_bulk_parallel, params: {prefix: users}}

Parameters:

Name Type Description Default
entity_list DataFrame

Entities whose mapped id column holds the MDM user ids to fetch.

required

assign_roles 🔗

assign_roles(user_id: str, role_ids: List[str]) -> Tuple[bool, str]

Assign roles to a user — additive (PATCH only, never revokes).

Use :meth:sync_roles when the supplied list should become the user's complete role set.

get_user_roles 🔗

get_user_roles(user_id: str) -> List[str]

Return the ids of the roles currently assigned to a user.

A failed lookup is reported as no roles, which is the long-standing contract here. Use :meth:_fetch_role_ids when the difference matters.

delete_role 🔗

delete_role(user_id: str, role_id: str) -> bool

Remove a single role from a user. A 404 counts as success (already gone).

sync_roles 🔗

sync_roles(user_id: str, new_role_ids: List[str], verify: bool = True) -> Tuple[bool, str]

Make new_role_ids the user's complete role set — destructive.

Uses PUT /users/{id}/roles, which the MDM API documents as "associate with the given roles and remove other existing associations" — a server-side replace in a single atomic call. The earlier read-then-DELETE-each-then-PATCH approach took 1+N+1 calls and left the user with no roles at all if it failed part-way; PUT cannot.

Unlike :meth:assign_roles this revokes roles the caller left out, which is what a back-office modification CSV means when a role is deleted from the Role cell.

Reads the current set first only to report what changed, and to skip the write entirely when the sets already match — so a bulk modification run doesn't churn every user's roles needlessly.

Note role ids are the role codes (verified against the live API: id == code for every role), so no code->id lookup is needed — the values in product_profiles.yml can be sent directly.

A 2xx does not mean the roles attached. The platform accepts the PUT and silently drops any role the customer cannot grant — its availableRoles, which are a property of the tenant's products. A Crop_intel analyst on a Digital_Ag-only tenant kept 2 of 3 roles while this reported "Roles added: APP_MONITORING|REGIONAL", and the account could not open the app. So the result is read back and compared, and a short apply is a failure with the missing roles named.

Parameters:

Name Type Description Default
user_id str

MDM user id.

required
new_role_ids List[str]

The complete desired role set.

required
verify bool

Read the roles back after the write and fail when any are missing. Costs one GET. Turn it off only for a bulk run that reconciles separately — the silent partial apply is the whole reason this check exists.

True

Returns:

Type Description
bool

(success, message) — the message names what was added/removed, or

str

what failed to attach, for surfacing in the run's Details column.

link_agronomist_to_growers(agronomist_id: str, grower_ids: List[str]) -> List[Dict]

Create manage links from AGRONOMIST to GROWERs.

link_grower_to_agronomist(grower_id: str, agronomist_id: str) -> Dict

Create managed-by link from GROWER to AGRONOMIST.

unlink_grower_from_agronomist(grower_id: str, agronomist_id: str) -> bool

Remove managed-by link from GROWER to AGRONOMIST.

unlink_agronomist_from_grower(agronomist_id: str, grower_id: str) -> bool

Remove manage link from AGRONOMIST to GROWER.

create_single_user 🔗

create_single_user(row: Series, send_login_details: bool = False, dry_run: bool = False) -> Tuple[bool, str, str]

Create a single user.

Parameters:

Name Type Description Default
send_login_details bool

Have the platform email the credentials.

False
dry_run bool

Log the call that would be made and change nothing. Returns a synthetic DRY_RUN:<login> id so the caller's later phases (roles, links) can be rehearsed against it too.

False

create_users 🔗

create_users(df: DataFrame, verbose: bool = True, send_login_details: bool = False, dry_run: bool = False) -> pd.DataFrame

Create multiple users with proper ordering (AGRONOMIST first, then GROWER).

⚠️ Reports only CREATED or ERROR — there is no "already exists" status, so re-running over the same input marks every row ERROR. Callers that need to resume should reconcile those against :meth:get_user_by_login.

Parameters:

Name Type Description Default
send_login_details bool

Have the platform email each created account its credentials. Off by default — this sends real email.

False
dry_run bool

Rehearse. Writes nothing, and logs every call all three phases would make — the user POST, the roles PATCH and the link PATCH — reporting status DRY_RUN.

It still performs the READS, which is what makes the rehearsal worth running: every login is checked with :meth:get_user_by_login, so the "no already-exists status" trap above surfaces here as WOULD COLLIDE instead of as a row of ERROR discovered after a live run has already created the other users. Link targets are resolved the same way, so an unresolvable Managed_by_user is reported before the write rather than silently skipped during it.

False

modify_users 🔗

modify_users(df: DataFrame, verbose: bool = True) -> pd.DataFrame

Modify multiple users in parallel. Synchronizes roles and managed-by links (adds new, removes old).

delete_users 🔗

delete_users(df: DataFrame, verbose: bool = True) -> pd.DataFrame

Delete multiple users in parallel.

export_credentials 🔗

export_credentials(results_df: DataFrame, df_users: DataFrame, output_path: str = None, separator: str = ';', statuses: Optional[List[str]] = None) -> Optional[str]

Write the credentials of newly created accounts to a CSV.

The generated password exists in exactly one place: memory. The create body carries it directly, the API never returns it, and the platform only emails it when sendLoginDetailsToUser was set. Re-run the cell that generated it and the value changes, so the account becomes untestable without a password reset — and a reset through :meth:modify_users must re-send email/login/userType or it blanks them. Persisting at creation time is the cheap way out of all of that.

Written to the git-ignored results/ tree, one timestamped file per run, alongside every other export. It is still a plaintext secret on disk: treat it as one, and delete it once the credentials are handed over.

Parameters:

Name Type Description Default
results_df DataFrame

Output of :meth:create_users, carrying Login / Status / User_Id.

required
df_users DataFrame

The input frame, which is where Password still lives — create_users does not echo it back.

required
output_path str

Defaults to the manager's results directory.

None
separator str

CSV separator, matching :meth:export_results.

';'
statuses Optional[List[str]]

Which rows to include. Defaults to ["CREATED"]; pass ["CREATED", "EXISTS"] on a resumed run, though EXISTS rows carry no usable password.

None

Returns:

Type Description
Optional[str]

str | None: Path written, or None when no row qualified.

export_results 🔗

export_results(results_df: DataFrame, operation: str, output_path: str = None, separator: str = ';') -> str

Export operation results to CSV.

get_latest_export 🔗

get_latest_export(output_path: str = None) -> Optional[str]

Get the most recent exported CSV file from results folder.