Native Asyncio¶
This section covers the native asyncio connection, cursors, and base classes.
Connection¶
DB API 2.0 interface to Amazon Athena: connect(), aio_connect(), and type objects.
- class pyathena.DBAPITypeObject[source]¶
A DB API type object that compares equal to each of its Athena type names.
https://www.python.org/dev/peps/pep-0249/#type-objects-and-constructors
- pyathena.connect(*args, cursor_class: None = ..., **kwargs) Connection[Cursor][source]¶
- pyathena.connect(*args, cursor_class: type[ConnectionCursor], **kwargs) Connection[ConnectionCursor]
Create a new database connection to Amazon Athena.
This function provides the main entry point for establishing connections to Amazon Athena. It follows the DB API 2.0 specification and returns a Connection object that can be used to create cursors for executing SQL queries.
- Parameters:
*args – Positional arguments passed to the Connection constructor, in the order of its parameters (
s3_staging_dir,region_name, …).s3_staging_dir – S3 location to store query results. Required if not using workgroups or if the workgroup doesn’t have a result location. Pass an empty string to explicitly disable S3 staging and skip the
AWS_ATHENA_S3_STAGING_DIRenvironment variable fallback (required for workgroups with managed query result storage).region_name – AWS region name. If not specified, uses the default region from your AWS configuration.
schema_name – Athena database/schema name. Defaults to “default”.
catalog_name – Athena data catalog name. Defaults to “awsdatacatalog”.
work_group – Athena workgroup name. Can be used instead of s3_staging_dir if the workgroup has a result location configured.
poll_interval – Time in seconds between polling for query completion. Defaults to 1.0.
encryption_option – S3 encryption option for query results. Can be “SSE_S3”, “SSE_KMS”, or “CSE_KMS”.
kms_key – KMS key ID for encryption when using SSE_KMS or CSE_KMS.
profile_name – AWS profile name to use for authentication.
role_arn – ARN of IAM role to assume for authentication.
role_session_name – Session name when assuming a role.
cursor_class – Custom cursor class to use. If not specified, uses the default Cursor class.
kill_on_interrupt – Whether to cancel running queries when interrupted. Defaults to True.
**kwargs – Additional keyword arguments passed to the Connection constructor.
- Returns:
A Connection object that can be used to create cursors and execute queries.
- Raises:
ProgrammingError – If neither s3_staging_dir nor work_group is provided.
Example
>>> import pyathena >>> conn = pyathena.connect( ... s3_staging_dir='s3://my-bucket/staging/', ... region_name='us-east-1', ... schema_name='mydatabase' ... ) >>> cursor = conn.cursor() >>> cursor.execute("SELECT * FROM mytable LIMIT 10") >>> results = cursor.fetchall()
- async pyathena.aio_connect(*args, **kwargs) AioConnection[source]¶
Create a new async database connection to Amazon Athena.
This is the async counterpart of
connect(). It returns anAioConnectionwhose cursors use nativeasynciofor polling and API calls, keeping the event loop free.- Parameters:
*args – Forwarded to
AioConnection.create(), which accepts keyword arguments only.**kwargs – Arguments forwarded to
AioConnection.create(). Seeconnect()for the full list of supported arguments.
- Returns:
An
AioConnectionthat producesAioCursorinstances by default.
Example
>>> import pyathena >>> conn = await pyathena.aio_connect( ... s3_staging_dir='s3://my-bucket/staging/', ... region_name='us-east-1', ... ) >>> async with conn.cursor() as cursor: ... await cursor.execute("SELECT 1") ... print(await cursor.fetchone())
- class pyathena.aio.connection.AioConnection(**kwargs: Any)[source]¶
Async-aware connection to Amazon Athena.
Wraps the synchronous
Connectionwith async context manager support and providescreate()for non-blocking initialization.Example
>>> async with await AioConnection.create( ... s3_staging_dir="s3://bucket/path/", ... region_name="us-east-1", ... ) as conn: ... async with conn.cursor() as cursor: ... await cursor.execute("SELECT 1") ... print(await cursor.fetchone())
- __init__(**kwargs: Any) None[source]¶
Initialize the connection with
AioCursoras the default cursor class.- Parameters:
**kwargs – Arguments forwarded to
Connection.__init__. If they do not includecursor_class, it is set toAioCursor.
- async classmethod create(**kwargs: Any) AioConnection[source]¶
Async factory for creating an
AioConnection.Runs the (potentially blocking)
__init__in a thread so that STS calls (role_arn/serial_number) do not block the loop.- Parameters:
**kwargs – Arguments forwarded to
AioConnection.__init__.- Returns:
A fully initialized
AioConnection.
- __enter__()¶
Enter the runtime context for the connection.
- Returns:
Self for use in context manager protocol.
- __exit__(exc_type, exc_val, exc_tb)¶
Exit the runtime context and close the connection.
- Parameters:
exc_type – Exception type if an exception occurred.
exc_val – Exception value if an exception occurred.
exc_tb – Exception traceback if an exception occurred.
- property client: BaseClient¶
Get the boto3 Athena client used for query operations.
- Returns:
The configured boto3 Athena client.
- close() None¶
Close the connection.
Closes the database connection. This method is provided for DB API 2.0 compatibility. Since Athena connections are stateless, this method currently does not perform any actual cleanup operations.
Note
This method is called automatically when using the connection as a context manager (with statement).
- commit() None¶
Commit any pending transaction.
This method is provided for DB API 2.0 compatibility. Since Athena does not support transactions, this method does nothing.
Note
Athena queries are auto-committed and cannot be rolled back.
- cursor(cursor: type[FunctionalCursor] | None = None, **kwargs) FunctionalCursor | ConnectionCursor¶
Create a new cursor object for executing queries.
Creates and returns a cursor object that can be used to execute SQL queries against Amazon Athena. The cursor inherits connection settings but can be customized with additional parameters.
- Parameters:
cursor – Custom cursor class to use. If not provided, uses the connection’s default cursor class.
**kwargs – Additional keyword arguments to pass to the cursor constructor. These override the connection’s
cursor_kwargsand its defaults.
- Returns:
A cursor object that can execute SQL queries.
Example
>>> cursor = connection.cursor() >>> cursor.execute("SELECT * FROM my_table LIMIT 10") >>> results = cursor.fetchall()
# Using a custom cursor type >>> from pyathena.pandas.cursor import PandasCursor >>> pandas_cursor = connection.cursor(PandasCursor) >>> df = pandas_cursor.execute(“SELECT * FROM my_table”).fetchall()
- property retry_config: RetryConfig¶
Get the retry configuration for AWS API calls.
- Returns:
The RetryConfig object that controls retry behavior for failed requests.
- rollback() None¶
Rollback any pending transaction.
This method is required by DB API 2.0 but is not supported by Athena since Athena does not support transactions.
- Raises:
NotSupportedError – Always raised since transactions are not supported.
- property session: Session¶
Get the boto3 session used for AWS API calls.
- Returns:
The configured boto3 Session object.
Aio Cursors¶
- class pyathena.aio.cursor.AioCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs)[source]¶
Native asyncio cursor for Amazon Athena.
Unlike
AsyncCursor(which usesThreadPoolExecutor), this cursor usesasyncio.sleepfor polling andasyncio.to_threadfor boto3 calls, keeping the event loop free.Example
>>> async with await AioConnection.create(...) as conn: ... async with conn.cursor() as cursor: ... await cursor.execute("SELECT * FROM my_table") ... rows = await cursor.fetchall()
- __init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs) None[source]¶
Initialize an AioCursor.
- Parameters:
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).
kms_key – KMS key for encryption.
kill_on_interrupt – Cancel the query when the task is cancelled while
execute()starts or waits for the query.result_reuse_enable – Enable Athena query result reuse.
result_reuse_minutes – Maximum age in minutes of a reused result.
**kwargs – Arguments forwarded to
WithResultSet.__init__andAioBaseCursor.__init__, such asarraysize,connection,converter,formatter, andretry_config.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) AioCursor[source]¶
Execute a SQL query asynchronously.
- Parameters:
operation – SQL query string to execute.
parameters – Query parameters (optional).
work_group – Athena workgroup to use (optional).
s3_staging_dir – S3 location for query results (optional).
cache_size – Number of queries to check for result caching (optional).
cache_expiration_time – Cache expiration time in seconds (optional).
result_reuse_enable – Enable result reuse (optional).
result_reuse_minutes – Result reuse duration in minutes (optional).
paramstyle – Parameter style to use (optional).
on_start_query_execution – Callback invoked with the query ID before
execute()waits for the query: after theStartQueryExecutioncall, or after a reusable query ID is found throughcache_size.result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
options – Shared execution options as an
ExecuteOptionsinstance. Individual keyword arguments take precedence overoptionsfields.**kwargs – Additional execution parameters.
- Returns:
Self reference for method chaining.
- async fetchone() Any | dict[Any, Any | None] | None[source]¶
Fetch the next row of a query result set.
- Returns:
The next row (a tuple, or a dict for
AioDictCursor), or None if no more rows.- Raises:
ProgrammingError – If called before executing a query that returns results.
- async fetchmany(size: int | None = None) list[Any | dict[Any, Any | None]][source]¶
Fetch multiple rows from a query result set.
- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
The fetched rows.
- Raises:
ProgrammingError – If called before executing a query that returns results.
- async fetchall() list[Any | dict[Any, Any | None]][source]¶
Fetch all remaining rows from a query result set.
- Returns:
The remaining rows.
- Raises:
ProgrammingError – If called before executing a query that returns results.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- async cancel() None¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- static get_default_converter(unload: bool = False) DefaultTypeConverter | Any¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
- class pyathena.aio.cursor.AioDictCursor(dict_type: type[Any] | None = None, **kwargs)[source]¶
Native asyncio cursor that returns rows as dictionaries.
Example
>>> async with await AioConnection.create(...) as conn: ... cursor = conn.cursor(AioDictCursor) ... await cursor.execute("SELECT id, name FROM users") ... row = await cursor.fetchone() ... print(row["name"])
- __init__(dict_type: type[Any] | None = None, **kwargs) None[source]¶
Initialize an AioDictCursor.
- Parameters:
dict_type – The type used to build each row of this cursor’s result sets. If None, the result set class’s
dict_typeis used.**kwargs – Arguments forwarded to
AioCursor.__init__.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- async cancel() None¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) AioCursor¶
Execute a SQL query asynchronously.
- Parameters:
operation – SQL query string to execute.
parameters – Query parameters (optional).
work_group – Athena workgroup to use (optional).
s3_staging_dir – S3 location for query results (optional).
cache_size – Number of queries to check for result caching (optional).
cache_expiration_time – Cache expiration time in seconds (optional).
result_reuse_enable – Enable result reuse (optional).
result_reuse_minutes – Result reuse duration in minutes (optional).
paramstyle – Parameter style to use (optional).
on_start_query_execution – Callback invoked with the query ID before
execute()waits for the query: after theStartQueryExecutioncall, or after a reusable query ID is found throughcache_size.result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
options – Shared execution options as an
ExecuteOptionsinstance. Individual keyword arguments take precedence overoptionsfields.**kwargs – Additional execution parameters.
- Returns:
Self reference for method chaining.
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- async fetchall() list[Any | dict[Any, Any | None]]¶
Fetch all remaining rows from a query result set.
- Returns:
The remaining rows.
- Raises:
ProgrammingError – If called before executing a query that returns results.
- async fetchmany(size: int | None = None) list[Any | dict[Any, Any | None]]¶
Fetch multiple rows from a query result set.
- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
The fetched rows.
- Raises:
ProgrammingError – If called before executing a query that returns results.
- async fetchone() Any | dict[Any, Any | None] | None¶
Fetch the next row of a query result set.
- Returns:
The next row (a tuple, or a dict for
AioDictCursor), or None if no more rows.- Raises:
ProgrammingError – If called before executing a query that returns results.
- static get_default_converter(unload: bool = False) DefaultTypeConverter | Any¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio Result Set¶
- class pyathena.aio.result_set.AthenaAioResultSet(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None)[source]¶
Async result set that provides async fetch methods.
Skips the synchronous
_pre_fetchby passing_pre_fetch=Falseto the parent__init__and provides anasync create()classmethod factory instead. Synchronous iteration raisesTypeError; useasync forinstead.- __init__(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None) None[source]¶
Initialize the result set without fetching rows;
create()fetches the first page.- Parameters:
connection – The connection that ran the query.
converter – The converter for result values.
query_execution – The query execution whose results to read.
arraysize – The number of rows per
GetQueryResultspage and the defaultfetchmany()size.retry_config – The retry configuration for API calls.
result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
- Raises:
ProgrammingError – If
query_executionis not given.
- async classmethod create(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None, **kwargs: Any) AthenaAioResultSet[source]¶
Async factory method.
Creates an
AthenaAioResultSetand awaits the initial data fetch.- Parameters:
connection – The database connection.
converter – Type converter for result values.
query_execution – Query execution metadata.
arraysize – Number of rows to fetch per request.
retry_config – Retry configuration for API calls.
result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
**kwargs – Additional arguments passed to the constructor of
cls, such asdict_typeforAthenaAioDictResultSet.
- Returns:
A fully initialized
AthenaAioResultSet.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None[source]¶
Fetch the next row of the result set.
Automatically fetches the next page from Athena when the current page is exhausted and more pages are available.
- Returns:
The next row (a tuple, or a dict for
AthenaAioDictResultSet), or None if no more rows.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]][source]¶
Fetch multiple rows from the result set.
- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
The rows, fewer than
sizewhen the result is exhausted.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]][source]¶
Fetch all remaining rows from the result set.
- Returns:
The remaining rows.
- __iter__() NoReturn[source]¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- property connection: Connection[Any]¶
The connection of the result set; raises
ProgrammingErrorif closed.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions.
None without result metadata, or for
INSERT,UPDATE,DELETE, andMERGE.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- property is_unload: bool¶
Check if the query is an UNLOAD statement.
- Returns:
True if the query is an UNLOAD statement, False otherwise.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
- class pyathena.aio.result_set.AthenaAioDictResultSet(*args: Any, dict_type: type[Any] | None = None, **kwargs: Any)[source]¶
Async result set that returns rows as dictionaries.
Inherits
_get_rowsfromAthenaDictResultSetand async fetch methods fromAthenaAioResultSetvia multiple inheritance.- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- __init__(*args: Any, dict_type: type[Any] | None = None, **kwargs: Any) None¶
Initialize the result set with an optional row type for this instance.
- Parameters:
*args – Positional arguments passed to the next
__init__in the MRO.dict_type – The type used to build each row of this result set. If None, the class attribute
dict_typeis used.**kwargs – Keyword arguments passed to the next
__init__in the MRO.
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- property connection: Connection[Any]¶
The connection of the result set; raises
ProgrammingErrorif closed.
- async classmethod create(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None, **kwargs: Any) AthenaAioResultSet¶
Async factory method.
Creates an
AthenaAioResultSetand awaits the initial data fetch.- Parameters:
connection – The database connection.
converter – Type converter for result values.
query_execution – Query execution metadata.
arraysize – Number of rows to fetch per request.
retry_config – Retry configuration for API calls.
result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
**kwargs – Additional arguments passed to the constructor of
cls, such asdict_typeforAthenaAioDictResultSet.
- Returns:
A fully initialized
AthenaAioResultSet.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions.
None without result metadata, or for
INSERT,UPDATE,DELETE, andMERGE.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch all remaining rows from the result set.
- Returns:
The remaining rows.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch multiple rows from the result set.
- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
The rows, fewer than
sizewhen the result is exhausted.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None¶
Fetch the next row of the result set.
Automatically fetches the next page from Athena when the current page is exhausted and more pages are available.
- Returns:
The next row (a tuple, or a dict for
AthenaAioDictResultSet), or None if no more rows.
- property is_unload: bool¶
Check if the query is an UNLOAD statement.
- Returns:
True if the query is an UNLOAD statement, False otherwise.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio Base Classes¶
- class pyathena.aio.common.AioBaseCursor(connection: Connection[Any], converter: Converter, formatter: Formatter, retry_config: RetryConfig, s3_staging_dir: str | None, schema_name: str | None, catalog_name: str | None, work_group: str | None, poll_interval: float, encryption_option: str | None, kms_key: str | None, kill_on_interrupt: bool, result_reuse_enable: bool, result_reuse_minutes: int, on_start_query_execution: Callable[[str], None] | None = None, on_poll: OnPollCallback | None = None, **kwargs)[source]¶
Async base cursor that overrides I/O methods with async equivalents.
Reuses
BaseCursor.__init__, all_build_*methods, and constants. Only the methods that perform network I/O or blocking sleep are overridden to useasyncio.to_thread/asyncio.sleep.- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase][source]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata[source]¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata][source]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __init__(connection: Connection[Any], converter: Converter, formatter: Formatter, retry_config: RetryConfig, s3_staging_dir: str | None, schema_name: str | None, catalog_name: str | None, work_group: str | None, poll_interval: float, encryption_option: str | None, kms_key: str | None, kill_on_interrupt: bool, result_reuse_enable: bool, result_reuse_minutes: int, on_start_query_execution: Callable[[str], None] | None = None, on_poll: OnPollCallback | None = None, **kwargs) None¶
Initialize the cursor with the settings it uses to run queries.
- Parameters:
connection – The connection that created the cursor.
converter – Converter for result values.
formatter – Formatter for query parameters.
retry_config – Retry configuration for API calls.
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).
kms_key – KMS key for encryption.
kill_on_interrupt – Cancel the execution when a
KeyboardInterruptinterrupts starting it or waiting for it.result_reuse_enable – Enable Athena query result reuse.
result_reuse_minutes – Maximum age in minutes of a reused result.
on_start_query_execution – Callback invoked with each query ID before the cursor waits for the query, by cursors whose
execute()supports it.on_poll – Callback invoked once per poll iteration with the current execution object.
**kwargs – Ignored.
- property connection: Connection[Any]¶
The connection that created this cursor.
- abstractmethod execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, **kwargs)¶
Execute a SQL query.
- Parameters:
operation – SQL query string.
parameters – Query parameters.
**kwargs – Execution options defined by the cursor implementation.
- abstractmethod executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query once for each set of parameters.
- Parameters:
operation – SQL query string.
seq_of_parameters – Sequence of parameter sets.
**kwargs – Execution options defined by the cursor implementation.
- static get_default_converter(unload: bool = False) DefaultTypeConverter | Any¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- class pyathena.aio.common.WithAsyncFetch(arraysize: int | None = None, **kwargs)[source]¶
Base class of the asyncio SQL cursors.
Combines
WithResultSetwithAioBaseCursorandCursorIterator, and provides async fetch,executemany, andcancel, async iteration, and the async context manager protocol. The fetch methods run the result set’s synchronous fetch withasyncio.to_thread; a subclass whose result set fetches asynchronously overrides them. Synchronous iteration raisesTypeError.Subclasses override
execute()and optionally__init__and format-specific helpers.- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None[source]¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- async cancel() None[source]¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None[source]¶
Fetch the next row of the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
A tuple representing the next row, or None if no more rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]][source]¶
Fetch multiple rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
List of tuples representing the fetched rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]][source]¶
Fetch all remaining rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
List of tuples representing all remaining rows.
- Raises:
ProgrammingError – If no result set is available.
- __iter__() NoReturn[source]¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __init__(arraysize: int | None = None, **kwargs) None¶
Initialize the cursor with no query ID and no result set.
- Parameters:
arraysize – Default number of rows per
fetchmany()call, validated by thearraysizesetter. If None,DEFAULT_FETCH_SIZEis used.**kwargs – Arguments passed to the next
__init__in the MRO.
- Raises:
ProgrammingError – If
arraysizeis outside the range the cursor’sarraysizesetter accepts.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- abstractmethod execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, **kwargs)¶
Execute a SQL query.
- Parameters:
operation – SQL query string.
parameters – Query parameters.
**kwargs – Execution options defined by the cursor implementation.
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- static get_default_converter(unload: bool = False) DefaultTypeConverter | Any¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio Pandas Cursor¶
- class pyathena.aio.pandas.cursor.AioPandasCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, engine: str = 'auto', chunksize: int | None = None, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, auto_optimize_chunksize: bool = False, **kwargs)[source]¶
Native asyncio cursor that returns results as pandas DataFrames.
Uses
asyncio.to_thread()for both result set creation and fetch operations, keeping the event loop free. This is especially important whenchunksizeis set, as fetch calls trigger lazy S3 reads.Example
>>> async with await pyathena.aio_connect(...) as conn: ... cursor = conn.cursor(AioPandasCursor) ... await cursor.execute("SELECT * FROM my_table") ... df = cursor.as_pandas()
- __init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, engine: str = 'auto', chunksize: int | None = None, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, auto_optimize_chunksize: bool = False, **kwargs) None[source]¶
Initialize an AioPandasCursor.
- Parameters:
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option for query results.
kms_key – KMS key for encrypting query results.
kill_on_interrupt – Cancel the query when the task is cancelled while
execute()starts or waits for the query.unload – Whether to wrap queries in
UNLOADand read the Parquet output.engine – Parsing engine (
auto,c,python, orpyarrow).chunksize – Number of rows per DataFrame chunk when reading CSV results. If set, it takes precedence over
auto_optimize_chunksize.block_size – Default block size of the S3 filesystem that reads the results.
cache_type – Default cache type of the S3 filesystem that reads the results.
max_workers – Maximum number of workers of the S3 filesystem.
result_reuse_enable – Whether to enable Athena query result reuse.
result_reuse_minutes – Maximum age of a reused query result in minutes.
auto_optimize_chunksize – Whether to choose a chunk size from the size of the CSV result file when
chunksizeis None.**kwargs – Other cursor arguments, such as
connectionandarraysize, passed to the parent__init__.
- static get_default_converter(unload: bool = False) DefaultPandasTypeConverter | Any[source]¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, keep_default_na: bool = False, na_values: Iterable[str] | None = ('',), quoting: int = 1, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) AioPandasCursor[source]¶
Execute a SQL query asynchronously and return results as pandas DataFrames.
- Parameters:
operation – SQL query string to execute.
parameters – Query parameters for parameterized queries.
work_group – Athena workgroup to use for this query.
s3_staging_dir – S3 location for query results.
cache_size – Number of queries to check for result caching.
cache_expiration_time – Cache expiration time in seconds.
result_reuse_enable – Enable Athena result reuse for this query.
result_reuse_minutes – Minutes to reuse cached results.
paramstyle – Parameter style (‘qmark’ or ‘pyformat’).
keep_default_na – Whether to keep default pandas NA values.
na_values – Additional values to treat as NA.
quoting – CSV quoting behavior (pandas csv.QUOTE_* constants).
on_start_query_execution – Callback invoked with the query ID before
execute()waits for the query: after theStartQueryExecutioncall, or after a reusable query ID is found throughcache_size.result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
options – Shared execution options as an
ExecuteOptionsinstance. Individual keyword arguments take precedence overoptionsfields.**kwargs – Additional pandas read_csv/read_parquet parameters.
- Returns:
Self reference for method chaining.
- as_pandas() DataFrame | PandasDataFrameIterator[source]¶
Return DataFrame or PandasDataFrameIterator based on chunksize setting.
- Returns:
DataFrame when chunksize is None, PandasDataFrameIterator when chunksize is set.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- async cancel() None¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch all remaining rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
List of tuples representing all remaining rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch multiple rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
List of tuples representing the fetched rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None¶
Fetch the next row of the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
A tuple representing the next row, or None if no more rows.
- Raises:
ProgrammingError – If no result set is available.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio Arrow Cursor¶
- class pyathena.aio.arrow.cursor.AioArrowCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, connect_timeout: float | None = None, request_timeout: float | None = None, **kwargs)[source]¶
Native asyncio cursor that returns results as Apache Arrow Tables.
Uses
asyncio.to_thread()for both result set creation and fetch operations, keeping the event loop free.Example
>>> async with await pyathena.aio_connect(...) as conn: ... cursor = conn.cursor(AioArrowCursor) ... await cursor.execute("SELECT * FROM my_table") ... table = cursor.as_arrow()
- __init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, connect_timeout: float | None = None, request_timeout: float | None = None, **kwargs) None[source]¶
Initialize an AioArrowCursor.
- Parameters:
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option for query results.
kms_key – KMS key for encrypting query results.
kill_on_interrupt – Cancel the query when the task is cancelled while
execute()starts or waits for the query.unload – Whether to wrap queries in
UNLOADand read the Parquet output.result_reuse_enable – Whether to enable Athena query result reuse.
result_reuse_minutes – Maximum age of a reused query result in minutes.
connect_timeout – Connection timeout in seconds of the pyarrow S3 filesystem that reads the results. If None, the pyarrow default is used.
request_timeout – Request timeout in seconds of the pyarrow S3 filesystem that reads the results. If None, the pyarrow default is used.
**kwargs – Other cursor arguments, such as
connectionandarraysize, passed to the parent__init__.
- static get_default_converter(unload: bool = False) DefaultArrowTypeConverter | DefaultArrowUnloadTypeConverter | Any[source]¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) AioArrowCursor[source]¶
Execute a SQL query asynchronously and return results as Arrow Tables.
- Parameters:
operation – SQL query string to execute.
parameters – Query parameters for parameterized queries.
work_group – Athena workgroup to use for this query.
s3_staging_dir – S3 location for query results.
cache_size – Number of queries to check for result caching.
cache_expiration_time – Cache expiration time in seconds.
result_reuse_enable – Enable Athena result reuse for this query.
result_reuse_minutes – Minutes to reuse cached results.
paramstyle – Parameter style (‘qmark’ or ‘pyformat’).
on_start_query_execution – Callback invoked with the query ID before
execute()waits for the query: after theStartQueryExecutioncall, or after a reusable query ID is found throughcache_size.result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
options – Shared execution options as an
ExecuteOptionsinstance. Individual keyword arguments take precedence overoptionsfields.**kwargs – Additional execution parameters.
- Returns:
Self reference for method chaining.
- as_arrow() Table[source]¶
Return query results as an Apache Arrow Table.
- Returns:
Apache Arrow Table containing all query results.
- as_polars() pl.DataFrame[source]¶
Return query results as a Polars DataFrame.
- Returns:
Polars DataFrame containing all query results.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- async cancel() None¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch all remaining rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
List of tuples representing all remaining rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch multiple rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
List of tuples representing the fetched rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None¶
Fetch the next row of the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
A tuple representing the next row, or None if no more rows.
- Raises:
ProgrammingError – If no result set is available.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio Polars Cursor¶
- class pyathena.aio.polars.cursor.AioPolarsCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, chunksize: int | None = None, **kwargs)[source]¶
Native asyncio cursor that returns results as Polars DataFrames.
Uses
asyncio.to_thread()for both result set creation and fetch operations, keeping the event loop free. This is especially important whenchunksizeis set, as fetch calls trigger lazy S3 reads.Example
>>> async with await pyathena.aio_connect(...) as conn: ... cursor = conn.cursor(AioPolarsCursor) ... await cursor.execute("SELECT * FROM my_table") ... df = cursor.as_polars()
- __init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, chunksize: int | None = None, **kwargs) None[source]¶
Initialize an AioPolarsCursor.
- Parameters:
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option for query results.
kms_key – KMS key for encrypting query results.
kill_on_interrupt – Cancel the query when the task is cancelled while
execute()starts or waits for the query.unload – Whether to wrap queries in
UNLOADand read the Parquet output.result_reuse_enable – Whether to enable Athena query result reuse.
result_reuse_minutes – Maximum age of a reused query result in minutes.
block_size – Default block size of the S3 filesystem that reads the results.
cache_type – Default cache type of the S3 filesystem that reads the results.
max_workers – Maximum number of workers of the S3 filesystem.
chunksize – Number of rows per chunk. If set, result files in S3 are read lazily in chunks of this size.
**kwargs – Other cursor arguments, such as
connectionandarraysize, passed to the parent__init__.
- static get_default_converter(unload: bool = False) DefaultPolarsTypeConverter | DefaultPolarsUnloadTypeConverter | Any[source]¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) AioPolarsCursor[source]¶
Execute a SQL query asynchronously and return results as Polars DataFrames.
- Parameters:
operation – SQL query string to execute.
parameters – Query parameters for parameterized queries.
work_group – Athena workgroup to use for this query.
s3_staging_dir – S3 location for query results.
cache_size – Number of queries to check for result caching.
cache_expiration_time – Cache expiration time in seconds.
result_reuse_enable – Enable Athena result reuse for this query.
result_reuse_minutes – Minutes to reuse cached results.
paramstyle – Parameter style (‘qmark’ or ‘pyformat’).
on_start_query_execution – Callback invoked with the query ID before
execute()waits for the query: after theStartQueryExecutioncall, or after a reusable query ID is found throughcache_size.result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
options – Shared execution options as an
ExecuteOptionsinstance. Individual keyword arguments take precedence overoptionsfields.**kwargs – Additional execution parameters passed to Polars read functions.
- Returns:
Self reference for method chaining.
- as_polars() pl.DataFrame[source]¶
Return query results as a Polars DataFrame.
- Returns:
Polars DataFrame containing all query results.
- as_arrow() Table[source]¶
Return query results as an Apache Arrow Table.
- Returns:
Apache Arrow Table containing all query results.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- async cancel() None¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch all remaining rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
List of tuples representing all remaining rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch multiple rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
List of tuples representing the fetched rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None¶
Fetch the next row of the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
A tuple representing the next row, or None if no more rows.
- Raises:
ProgrammingError – If no result set is available.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio S3FS Cursor¶
- class pyathena.aio.s3fs.cursor.AioS3FSCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, csv_reader: type[DefaultCSVReader] | type[AthenaCSVReader] | None = None, **kwargs)[source]¶
Native asyncio cursor that reads CSV results via AioS3FileSystem.
Uses
AioS3FileSystemfor S3 operations, which replacesThreadPoolExecutorparallelism withasyncio.gather+asyncio.to_thread. Fetch operations are wrapped inasyncio.to_thread()because CSV reading is blocking I/O.Example
>>> async with await pyathena.aio_connect(...) as conn: ... cursor = conn.cursor(AioS3FSCursor) ... await cursor.execute("SELECT * FROM my_table") ... row = await cursor.fetchone()
- __init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, csv_reader: type[DefaultCSVReader] | type[AthenaCSVReader] | None = None, **kwargs) None[source]¶
Initialize an AioS3FSCursor.
- Parameters:
s3_staging_dir – S3 location for query results.
schema_name – Default schema name.
catalog_name – Default catalog name.
work_group – Athena workgroup name.
poll_interval – Query status polling interval in seconds.
encryption_option – S3 encryption option for query results.
kms_key – KMS key for encrypting query results.
kill_on_interrupt – Cancel the query when the task is cancelled while
execute()starts or waits for the query.result_reuse_enable – Whether to enable Athena query result reuse.
result_reuse_minutes – Maximum age of a reused query result in minutes.
csv_reader – CSV reader class for parsing the result files. If None,
AthenaCSVReaderis used, which distinguishes NULL from empty strings.DefaultCSVReaderreads both as empty strings.**kwargs – Other cursor arguments, such as
connectionandarraysize, passed to the parent__init__.
- static get_default_converter(unload: bool = False) DefaultS3FSTypeConverter[source]¶
Get the default type converter for S3FS cursor.
- Parameters:
unload – Unused. S3FS cursor does not support UNLOAD operations.
- Returns:
DefaultS3FSTypeConverter instance.
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) AioS3FSCursor[source]¶
Execute a SQL query asynchronously via S3FileSystem CSV reader.
- Parameters:
operation – SQL query string to execute.
parameters – Query parameters for parameterized queries.
work_group – Athena workgroup to use for this query.
s3_staging_dir – S3 location for query results.
cache_size – Number of queries to check for result caching.
cache_expiration_time – Cache expiration time in seconds.
result_reuse_enable – Enable Athena result reuse for this query.
result_reuse_minutes – Minutes to reuse cached results.
paramstyle – Parameter style (‘qmark’ or ‘pyformat’).
on_start_query_execution – Callback invoked with the query ID before
execute()waits for the query: after theStartQueryExecutioncall, or after a reusable query ID is found throughcache_size.result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.
options – Shared execution options as an
ExecuteOptionsinstance. Individual keyword arguments take precedence overoptionsfields.**kwargs – Additional execution parameters.
- Returns:
Self reference for method chaining.
- DEFAULT_RESULT_REUSE_MINUTES = 60¶
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __iter__() NoReturn¶
Reject synchronous iteration; use
async forinstead.- Raises:
TypeError – Always, because the fetch methods are coroutines.
- property arraysize: int¶
The default number of rows per
fetchmany()call.execute()passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raisesProgrammingError.- Returns:
The default number of rows per
fetchmany()call.
- async cancel() None¶
Cancel the currently executing query.
- Raises:
ProgrammingError – If no query is currently executing.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property data_manifest_location: str | None¶
The S3 location of the data manifest that lists the files the query wrote.
- property description: list[tuple[str, str, None, None, int, int, str]] | None¶
The DB API 2.0 column descriptions of the result set, or None without one.
- property encryption_option: str | None¶
The
EncryptionOptionof the query results, such asSSE_S3orSSE_KMS.
- property engine_execution_time_in_millis: int | None¶
The time in milliseconds that the query engine took to run the query.
- property error_category: int | None¶
1 for system, 2 for user, 3 for other.
- Type:
The
ErrorCategoryof the failure
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None¶
Execute a SQL query multiple times with different parameters.
On success,
rowcountis the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.On failure,
query_idretains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.- Parameters:
operation – SQL query string to execute.
seq_of_parameters – Sequence of parameter sets, one per execution.
**kwargs – Additional keyword arguments passed to each
execute().
- property expected_bucket_owner: str | None¶
The AWS account ID expected to own the S3 bucket of the query results.
- async fetchall() list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch all remaining rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
List of tuples representing all remaining rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchmany(size: int | None = None) list[tuple[Any | None, ...] | dict[Any, Any | None]]¶
Fetch multiple rows from the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Parameters:
size – Maximum number of rows to fetch. If None or not positive,
arraysizeis used.- Returns:
List of tuples representing the fetched rows.
- Raises:
ProgrammingError – If no result set is available.
- async fetchone() tuple[Any | None, ...] | dict[Any, Any | None] | None¶
Fetch the next row of the result set.
Wraps the synchronous fetch in
asyncio.to_threadto avoid blocking the event loop.- Returns:
A tuple representing the next row, or None if no more rows.
- Raises:
ProgrammingError – If no result set is available.
- async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- async list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- property query_id: str | None¶
The query execution ID of the last execution.
With
cache_sizeorcache_expiration_time, this can be the ID of a previous execution whose result is reused.- Returns:
The query execution ID, or None if there is none since the last reset.
- property query_planning_time_in_millis: int | None¶
The time in milliseconds that Athena took to plan the query.
- property query_queue_time_in_millis: int | None¶
The time in milliseconds that the query waited in the queue.
- property result_reuse_enabled: bool | None¶
Whether reuse of previous query results by age is enabled for the query.
- property result_reuse_minutes: int | None¶
The maximum age in minutes of a previous query result that Athena can reuse.
- property result_set: AthenaResultSet | None¶
The result set of the last executed query.
- Returns:
The result set, or None before a query succeeds or after a reset.
- property reused_previous_result: bool | None¶
Whether Athena reused a previous query result instead of running the query.
- property rowcount: int¶
Get the number of rows affected by the last operation.
For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful
executemany(), this is the sum across executions, or -1 if any count is unknown.- Returns:
The number of rows, or -1 if not applicable or unknown.
- property rownumber: int | None¶
The zero-based index of the next row in the result set.
- Returns:
The row index, or None if there is no result set or the index is unknown.
- property s3_acl_option: str | None¶
The
S3AclOptionof the query results, such asBUCKET_OWNER_FULL_CONTROL.
- property service_processing_time_in_millis: int | None¶
The time in milliseconds that Athena took to publish the query results.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonthat gives further detail about the state.
Aio Spark Cursor¶
- class pyathena.aio.spark.cursor.AioSparkCursor(session_id: str | None = None, description: str | None = None, engine_configuration: dict[str, Any] | None = None, notebook_version: str | None = None, session_idle_timeout_minutes: int | None = None, terminate_session_on_close: bool | None = None, **kwargs)[source]¶
Native asyncio cursor for executing PySpark code on Athena.
Overrides post-init I/O methods of
SparkBaseCursorwith async equivalents. Session management (_exists_session,_start_session, etc.) stays synchronous because__init__runs insideasyncio.to_thread.Since
SparkBaseCursor.__init__performs I/O (session management), cursor creation must be wrapped inasyncio.to_thread:cursor = await asyncio.to_thread(conn.cursor)
Example
>>> import asyncio >>> async with await pyathena.aio_connect( ... work_group="spark-workgroup", ... cursor_class=AioSparkCursor, ... ) as conn: ... cursor = await asyncio.to_thread(conn.cursor) ... await cursor.execute("spark.sql('SELECT 1').show()") ... print(await cursor.get_std_out())
- property calculation_execution: AthenaCalculationExecution | None¶
The calculation execution that the other properties read, or None.
- async get_std_out() str | None[source]¶
Get the standard output from the Spark calculation execution.
- Returns:
The standard output as a string, or None if no output is available.
- async get_std_error() str | None[source]¶
Get the standard error from the Spark calculation execution.
- Returns:
The standard error as a string, or None if no error output is available.
- async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, session_id: str | None = None, description: str | None = None, client_request_token: str | None = None, work_group: str | None = None, **kwargs) AioSparkCursor[source]¶
Execute PySpark code asynchronously.
- Parameters:
operation – PySpark code to execute.
parameters – Unused, kept for API compatibility.
session_id – Spark session ID override.
description – Calculation description.
client_request_token – Idempotency token.
work_group – Unused, kept for API compatibility.
**kwargs – Additional parameters.
- Returns:
Self reference for method chaining.
- async cancel() None[source]¶
Cancel the currently running calculation.
- Raises:
ProgrammingError – If no calculation is running.
- async close() None[source]¶
Close the cursor, terminating its Spark session if configured to.
See
terminate_session_on_close. After a successful termination, further calls do not terminate the session again; after a failed one, calling this method again retries it.- Raises:
OperationalError – If terminating the session fails.
- async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) None[source]¶
Execute a SQL query once for each set of parameters.
- Parameters:
operation – SQL query string.
seq_of_parameters – Sequence of parameter sets.
**kwargs – Execution options defined by the cursor implementation.
- LIST_DATABASES_MAX_RESULTS = 50¶
- LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50¶
- LIST_TABLE_METADATA_MAX_RESULTS = 50¶
- __init__(session_id: str | None = None, description: str | None = None, engine_configuration: dict[str, Any] | None = None, notebook_version: str | None = None, session_idle_timeout_minutes: int | None = None, terminate_session_on_close: bool | None = None, **kwargs) None¶
Initialize the cursor and start or attach to a Spark session.
If waiting for a newly started session fails, that session is terminated regardless of
terminate_session_on_close; a supplied session is not.- Parameters:
session_id – ID of an existing session to use. If omitted, a new session is started.
description – Description of a new session.
engine_configuration – Engine configuration of a new session. Defaults to
get_default_engine_configuration().notebook_version – Notebook version of a new session.
session_idle_timeout_minutes – Idle timeout of a new session in minutes.
terminate_session_on_close – Whether
close()terminates the session. If None, only a session started by this cursor is terminated; a session supplied withsession_idis left running.**kwargs – Arguments passed to
BaseCursor.
- Raises:
OperationalError – If the supplied session does not exist, or the session cannot be started or does not become idle.
- property calculation_id: str | None¶
The ID of the calculation tracked by this cursor, or None if there is none.
- property completion_date_time: datetime | None¶
The
CompletionDateTimeof the calculation, or None if there is none.
- property connection: Connection[Any]¶
The connection that created this cursor.
- property dpu_execution_in_millis: int | None¶
The
DpuExecutionInMillisstatistic of the calculation, or None if there is none.
- static get_default_converter(unload: bool = False) DefaultTypeConverter | Any¶
Get the default type converter for this cursor class.
- Parameters:
unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.
- Returns:
The default type converter instance for this cursor type.
- static get_default_engine_configuration() dict[str, Any]¶
Return the engine configuration used when none is given.
- Returns:
a coordinator DPU size of 1, at most 2 concurrent DPUs, and a default executor DPU size of 1.
- Return type:
The
EngineConfigurationof a new session
- get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) AthenaTableMetadata¶
Get one table’s metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
table_name – The table name.
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
logging – Whether to log a failed request.
- Returns:
The table’s metadata.
- Raises:
OperationalError – If the request fails, including when the table does not exist.
- list_databases(catalog_name: str | None, max_results: int | None = None) list[AthenaDatabase]¶
List the catalog’s databases.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
max_results – The page size of each request.
- Returns:
The catalog’s databases.
- Raises:
OperationalError – If the request fails.
- list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) list[AthenaTableMetadata]¶
List a database’s table metadata.
In
AwsDataCatalogand S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; seeglue_metadata_fallback.- Parameters:
catalog_name – The catalog, or None for the cursor’s catalog.
schema_name – The database, or None for the cursor’s schema.
expression – A table name pattern.
max_results – The page size of each request.
logging – Whether to log a failed request.
- Returns:
The metadata of the database’s tables.
- Raises:
OperationalError – If the request fails.
- setinputsizes(sizes)¶
Accept input sizes as DB API 2.0 requires, and ignore them.
- Parameters:
sizes – Sequence of parameter types or sizes.
- setoutputsize(size, column=None)¶
Accept a column buffer size as DB API 2.0 requires, and ignore it.
- Parameters:
size – Buffer size for large columns.
column – Index of the column the size applies to, or None for all large columns.
- property state_change_reason: str | None¶
The
StateChangeReasonof the calculation, or None if there is none.
- property std_error_s3_uri: str | None¶
The
StdErrorS3Uriof the calculation, or None if there is none.