Utilities and Configuration¶
This section covers utility functions, retry configuration, and helper classes.
Retry Configuration¶
- class pyathena.util.RetryConfig(exceptions: Iterable[str] = ('ThrottlingException', 'TooManyRequestsException'), attempt: int = 8, multiplier: int = 1, max_delay: int = 100, exponential_base: int = 2)[source]¶
Configuration for automatic retry behavior on failed API calls.
This class configures how PyAthena handles transient failures when communicating with AWS services. It uses exponential backoff with customizable parameters to retry failed operations.
- exceptions¶
Tuple of AWS exception names to retry on.
- attempt¶
Maximum number of attempts, including the first call.
- multiplier¶
Base multiplier for exponential backoff in seconds. Each wait also adds uniform random jitter of up to one multiplier.
- max_delay¶
Maximum exponential delay between retries in seconds, before jitter is added.
- exponential_base¶
Base for exponential backoff calculation.
Example
>>> from pyathena.util import RetryConfig >>> >>> # Default retry configuration >>> retry_config = RetryConfig() >>> >>> # Custom retry configuration >>> custom_retry = RetryConfig( ... exceptions=["ThrottlingException", "ServiceUnavailableException"], ... attempt=10, ... max_delay=60 ... ) >>> >>> # Use with connection >>> conn = pyathena.connect( ... s3_staging_dir="s3://bucket/path/", ... retry_config=custom_retry ... )
Note
Exception names may be supplied as a single string or an iterable. They are captured as a tuple at construction; later changes to the original iterable do not change this configuration. Retries are applied to AWS API calls, not to SQL query execution. Query failures typically require manual intervention or query fixes. With the default settings, the exponential waits between attempts sum to 127 seconds plus jitter, which outlasts the metadata API throttling episodes observed under concurrent reflection. SDK retries configured on the boto3 client are a separate layer applied within each attempt. Recognized Glue error codes wrapped in Athena MetadataException are matched against exceptions in the same way as direct AWS error codes.
- __init__(exceptions: Iterable[str] = ('ThrottlingException', 'TooManyRequestsException'), attempt: int = 8, multiplier: int = 1, max_delay: int = 100, exponential_base: int = 2) None[source]¶
Initialize the retry configuration.
- Parameters:
exceptions – AWS error code, or iterable of error codes, to retry on. Stored as a tuple.
attempt – Maximum number of attempts, including the first call.
multiplier – Base multiplier for exponential backoff in seconds, and the upper bound of the random jitter added to each wait.
max_delay – Maximum exponential delay between retries in seconds, before jitter is added.
exponential_base – Base for exponential backoff calculation.
Utility Functions¶
- pyathena.util.retry_api_call(func: Callable[[...], Any], config: RetryConfig, logger: Logger | None = None, *args, stop_on: Callable[[BaseException], bool] | None = None, **kwargs) Any[source]¶
Execute a function with automatic retry logic for AWS API calls.
This function wraps AWS API calls with retry behavior based on the provided configuration. It uses exponential backoff with uniform jitter and only retries on specific AWS exceptions that indicate transient failures.
- Parameters:
func – The AWS API function to call.
config – RetryConfig instance specifying retry behavior.
logger – Optional logger for retry attempt logging.
*args – Positional arguments to pass to the function.
stop_on – Optional predicate; an exception it accepts is raised at once instead of being retried.
**kwargs – Keyword arguments to pass to the function.
- Returns:
The result of the successful function call.
- Raises:
The original exception if all retry attempts are exhausted. –
Example
>>> from pyathena.util import RetryConfig, retry_api_call >>> config = RetryConfig(attempt=3, max_delay=30) >>> result = retry_api_call( ... client.describe_table, ... config=config, ... logger=logger, ... TableName="my_table" ... )
Note
Only retries on AWS exceptions listed in the RetryConfig.exceptions. This includes recognized Glue error codes wrapped in MetadataException. Other errors are propagated without retrying.
- pyathena.util.is_retryable_error(ex: BaseException, config: RetryConfig) bool[source]¶
Return whether an exception matches the retry policy’s AWS error codes.
- Parameters:
ex – The exception raised by an AWS API call.
config – RetryConfig whose
exceptionslist the retryable error codes.
- Returns:
True if the direct error code, or a recognized Glue error code wrapped in an Athena
MetadataException, is listed inconfig.exceptions.
- pyathena.util.parse_output_location(output_location: str) tuple[str, str][source]¶
Parse an S3 output location URL into bucket and key components.
- Parameters:
output_location – S3 URL in format ‘s3://bucket-name/path/to/object’
- Returns:
Tuple of (bucket_name, object_key)
- Raises:
DataError – If the output_location format is invalid.
Example
>>> bucket, key = parse_output_location("s3://my-bucket/results/query.csv") >>> print(bucket) # "my-bucket" >>> print(key) # "results/query.csv"
- pyathena.util.strtobool(val)[source]¶
Convert a string representation of truth to True or False.
This function replaces the deprecated distutils.util.strtobool method. It converts string representations of boolean values to actual boolean values.
- Parameters:
val – String representation of a boolean value.
- Returns:
1 for True values, 0 for False values.
- Raises:
ValueError – If the input string is not a recognized boolean representation.
Example
>>> strtobool("yes") # 1 >>> strtobool("false") # 0 >>> strtobool("invalid") # ValueError
Note
True values: y, yes, t, true, on, 1 (case-insensitive) False values: n, no, f, false, off, 0 (case-insensitive)
References
Common Base Classes¶
- class pyathena.common.CursorIterator(**kwargs)[source]¶
Abstract base class providing iteration and result fetching capabilities for cursors.
This mixin class provides common functionality for iterating through query results and managing cursor state. It implements the iterator protocol and provides standard fetch methods that conform to the DB API 2.0 specification.
- DEFAULT_RESULT_REUSE_MINUTES¶
Default minutes for Athena result reuse (60).
- arraysize¶
Number of rows to fetch with fetchmany() if size not specified.
Note
This is an abstract base class used by concrete cursor implementations. It should not be instantiated directly.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- __init__(**kwargs) None[source]¶
Initialize the iterator with no current row and an unknown row count.
- Parameters:
**kwargs – Keyword arguments, of which only
arraysizeis used. If it is absent,DEFAULT_FETCH_SIZEis used.- Raises:
ProgrammingError – If
arraysizeis outside the range thearraysizesetter accepts.
- class pyathena.common.BaseCursor(connection: Connection[Any], converter: Converter, formatter: Formatter, retry_config: RetryConfig, s3_staging_dir: str | None, schema_name: str | None, catalog_name: str | None, work_group: str | None, poll_interval: float, encryption_option: str | None, kms_key: str | None, kill_on_interrupt: bool, result_reuse_enable: bool, result_reuse_minutes: int, on_start_query_execution: Callable[[str], None] | None = None, on_poll: OnPollCallback | None = None, **kwargs)[source]¶
Abstract base class for all PyAthena cursor implementations.
This class provides the foundational functionality for executing SQL queries and calculations on Amazon Athena. It handles AWS API interactions, query execution management, metadata operations, and result polling.
All concrete cursor implementations (Cursor, DictCursor, PandasCursor, ArrowCursor, SparkCursor, AsyncCursor) inherit from this base class and implement the abstract methods according to their specific use cases.
- LIST_QUERY_EXECUTIONS_MAX_RESULTS¶
Maximum results per query listing API call (50).
- LIST_TABLE_METADATA_MAX_RESULTS¶
Maximum results per table metadata API call (50).
- LIST_DATABASES_MAX_RESULTS¶
Maximum results per database listing API call (50).
- Key Features:
Query execution and polling with configurable retry logic
Table and database metadata operations
Result caching and reuse capabilities
Encryption and security configuration support
Workgroup and catalog management
Query cancellation and interruption handling
Example
This is an abstract base class and should not be instantiated directly. Use concrete implementations like Cursor or PandasCursor instead:
>>> cursor = connection.cursor() # Creates default Cursor >>> cursor.execute("SELECT * FROM my_table") >>> results = cursor.fetchall()
Note
This class contains AWS service quotas as constants. These limits are enforced by the AWS Athena service and should not be modified.
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- __init__(connection: Connection[Any], converter: Converter, formatter: Formatter, retry_config: RetryConfig, s3_staging_dir: str | None, schema_name: str | None, catalog_name: str | None, work_group: str | None, poll_interval: float, encryption_option: str | None, kms_key: str | None, kill_on_interrupt: bool, result_reuse_enable: bool, result_reuse_minutes: int, on_start_query_execution: Callable[[str], None] | None = None, on_poll: OnPollCallback | None = None, **kwargs) None[source]¶
Initialize the cursor with the settings it uses to run queries.
- Parameters:
connection – The connection that created the cursor.
converter – Converter for result values.
formatter – Formatter for query parameters.
retry_config – Retry configuration for API calls.
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).
kms_key – KMS key for encryption.
kill_on_interrupt – Cancel the execution when a
KeyboardInterruptinterrupts starting it or waiting for it.result_reuse_enable – Enable Athena query result reuse.
result_reuse_minutes – Maximum age in minutes of a reused result.
on_start_query_execution – Callback invoked with each query ID before the cursor waits for the query, by cursors whose
execute()supports it.on_poll – Callback invoked once per poll iteration with the current execution object.
**kwargs – Ignored.
- static get_default_converter(unload: bool = False) DefaultTypeConverter | Any[source]¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- property connection: Connection[Any]¶
The connection that created this cursor.
- list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase][source]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata[source]¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata][source]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- abstractmethod execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, **kwargs)[source]¶
Execute a SQL query.
- Parameters:
operation – SQL query string.
parameters – Query parameters.
**kwargs – Execution options defined by the cursor implementation.
- abstractmethod executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None[source]¶
Execute a SQL query once for each set of parameters.
- Parameters:
operation – SQL query string.
seq_of_parameters – Sequence of parameter sets.
**kwargs – Execution options defined by the cursor implementation.