Native Asyncio

This section covers the native asyncio connection, cursors, and base classes.

Connection

DB API 2.0 interface to Amazon Athena: connect(), aio_connect(), and type objects.

class pyathena.DBAPITypeObject[source]

A DB API type object that compares equal to each of its Athena type names.

https://www.python.org/dev/peps/pep-0249/#type-objects-and-constructors

pyathena.connect(*args, cursor_class: None = ..., **kwargs) → Connection[Cursor][source]
pyathena.connect(*args, cursor_class: type[ConnectionCursor], **kwargs) → Connection[ConnectionCursor]

Create a new database connection to Amazon Athena.

This function provides the main entry point for establishing connections to Amazon Athena. It follows the DB API 2.0 specification and returns a Connection object that can be used to create cursors for executing SQL queries.

Parameters:
  • *args – Positional arguments passed to the Connection constructor, in the order of its parameters (s3_staging_dir, region_name, …).

  • s3_staging_dir – S3 location to store query results. Required if not using workgroups or if the workgroup doesn’t have a result location. Pass an empty string to explicitly disable S3 staging and skip the AWS_ATHENA_S3_STAGING_DIR environment variable fallback (required for workgroups with managed query result storage).

  • region_name – AWS region name. If not specified, uses the default region from your AWS configuration.

  • schema_name – Athena database/schema name. Defaults to “default”.

  • catalog_name – Athena data catalog name. Defaults to “awsdatacatalog”.

  • work_group – Athena workgroup name. Can be used instead of s3_staging_dir if the workgroup has a result location configured.

  • poll_interval – Time in seconds between polling for query completion. Defaults to 1.0.

  • encryption_option – S3 encryption option for query results. Can be “SSE_S3”, “SSE_KMS”, or “CSE_KMS”.

  • kms_key – KMS key ID for encryption when using SSE_KMS or CSE_KMS.

  • profile_name – AWS profile name to use for authentication.

  • role_arn – ARN of IAM role to assume for authentication.

  • role_session_name – Session name when assuming a role.

  • cursor_class – Custom cursor class to use. If not specified, uses the default Cursor class.

  • kill_on_interrupt – Whether to cancel running queries when interrupted. Defaults to True.

  • **kwargs – Additional keyword arguments passed to the Connection constructor.

Returns:

A Connection object that can be used to create cursors and execute queries.

Raises:

ProgrammingError – If neither s3_staging_dir nor work_group is provided.

Example

>>> import pyathena
>>> conn = pyathena.connect(
...     s3_staging_dir='s3://my-bucket/staging/',
...     region_name='us-east-1',
...     schema_name='mydatabase'
... )
>>> cursor = conn.cursor()
>>> cursor.execute("SELECT * FROM mytable LIMIT 10")
>>> results = cursor.fetchall()
async pyathena.aio_connect(*args, **kwargs) → AioConnection[source]

Create a new async database connection to Amazon Athena.

This is the async counterpart of connect(). It returns an AioConnection whose cursors use native asyncio for polling and API calls, keeping the event loop free.

Parameters:
  • *args – Forwarded to AioConnection.create(), which accepts keyword arguments only.

  • **kwargs – Arguments forwarded to AioConnection.create(). See connect() for the full list of supported arguments.

Returns:

An AioConnection that produces AioCursor instances by default.

Example

>>> import pyathena
>>> conn = await pyathena.aio_connect(
...     s3_staging_dir='s3://my-bucket/staging/',
...     region_name='us-east-1',
... )
>>> async with conn.cursor() as cursor:
...     await cursor.execute("SELECT 1")
...     print(await cursor.fetchone())
class pyathena.aio.connection.AioConnection(**kwargs: Any)[source]

Async-aware connection to Amazon Athena.

Wraps the synchronous Connection with async context manager support and provides create() for non-blocking initialization.

Example

>>> async with await AioConnection.create(
...     s3_staging_dir="s3://bucket/path/",
...     region_name="us-east-1",
... ) as conn:
...     async with conn.cursor() as cursor:
...         await cursor.execute("SELECT 1")
...         print(await cursor.fetchone())
__init__(**kwargs: Any) → None[source]

Initialize the connection with AioCursor as the default cursor class.

Parameters:

**kwargs – Arguments forwarded to Connection.__init__. If they do not include cursor_class, it is set to AioCursor.

async classmethod create(**kwargs: Any) → AioConnection[source]

Async factory for creating an AioConnection.

Runs the (potentially blocking) __init__ in a thread so that STS calls (role_arn / serial_number) do not block the loop.

Parameters:

**kwargs – Arguments forwarded to AioConnection.__init__.

Returns:

A fully initialized AioConnection.

__enter__()

Enter the runtime context for the connection.

Returns:

Self for use in context manager protocol.

__exit__(exc_type, exc_val, exc_tb)

Exit the runtime context and close the connection.

Parameters:
  • exc_type – Exception type if an exception occurred.

  • exc_val – Exception value if an exception occurred.

  • exc_tb – Exception traceback if an exception occurred.

property client: BaseClient

Get the boto3 Athena client used for query operations.

Returns:

The configured boto3 Athena client.

close() → None

Close the connection.

Closes the database connection. This method is provided for DB API 2.0 compatibility. Since Athena connections are stateless, this method currently does not perform any actual cleanup operations.

Note

This method is called automatically when using the connection as a context manager (with statement).

commit() → None

Commit any pending transaction.

This method is provided for DB API 2.0 compatibility. Since Athena does not support transactions, this method does nothing.

Note

Athena queries are auto-committed and cannot be rolled back.

cursor(cursor: type[FunctionalCursor] | None = None, **kwargs) → FunctionalCursor | ConnectionCursor

Create a new cursor object for executing queries.

Creates and returns a cursor object that can be used to execute SQL queries against Amazon Athena. The cursor inherits connection settings but can be customized with additional parameters.

Parameters:
  • cursor – Custom cursor class to use. If not provided, uses the connection’s default cursor class.

  • **kwargs – Additional keyword arguments to pass to the cursor constructor. These override the connection’s cursor_kwargs and its defaults.

Returns:

A cursor object that can execute SQL queries.

Example

>>> cursor = connection.cursor()
>>> cursor.execute("SELECT * FROM my_table LIMIT 10")
>>> results = cursor.fetchall()

# Using a custom cursor type >>> from pyathena.pandas.cursor import PandasCursor >>> pandas_cursor = connection.cursor(PandasCursor) >>> df = pandas_cursor.execute(“SELECT * FROM my_table”).fetchall()

property retry_config: RetryConfig

Get the retry configuration for AWS API calls.

Returns:

The RetryConfig object that controls retry behavior for failed requests.

rollback() → None

Rollback any pending transaction.

This method is required by DB API 2.0 but is not supported by Athena since Athena does not support transactions.

Raises:

NotSupportedError – Always raised since transactions are not supported.

property session: Session

Get the boto3 session used for AWS API calls.

Returns:

The configured boto3 Session object.

Aio Cursors

class pyathena.aio.cursor.AioCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs)[source]

Native asyncio cursor for Amazon Athena.

Unlike AsyncCursor (which uses ThreadPoolExecutor), this cursor uses asyncio.sleep for polling and asyncio.to_thread for boto3 calls, keeping the event loop free.

Example

>>> async with await AioConnection.create(...) as conn:
...     async with conn.cursor() as cursor:
...         await cursor.execute("SELECT * FROM my_table")
...         rows = await cursor.fetchall()
__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs) → None[source]

Initialize an AioCursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).

  • kms_key – KMS key for encryption.

  • kill_on_interrupt – Cancel the query when the task is cancelled while execute() starts or waits for the query.

  • result_reuse_enable – Enable Athena query result reuse.

  • result_reuse_minutes – Maximum age in minutes of a reused result.

  • **kwargs – Arguments forwarded to WithResultSet.__init__ and AioBaseCursor.__init__, such as arraysize, connection, converter, formatter, and retry_config.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → AioCursor[source]

Execute a SQL query asynchronously.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters (optional).

  • work_group – Athena workgroup to use (optional).

  • s3_staging_dir – S3 location for query results (optional).

  • cache_size – Number of queries to check for result caching (optional).

  • cache_expiration_time – Cache expiration time in seconds (optional).

  • result_reuse_enable – Enable result reuse (optional).

  • result_reuse_minutes – Result reuse duration in minutes (optional).

  • paramstyle – Parameter style to use (optional).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

Self reference for method chaining.

async fetchone() → Any | dict[Any, Any | None] | None[source]

Fetch the next row of a query result set.

Returns:

The next row (a tuple, or a dict for AioDictCursor), or None if no more rows.

Raises:

ProgrammingError – If called before executing a query that returns results.

async fetchmany(size: int | None = None) → list[Any | dict[Any, Any | None]][source]

Fetch multiple rows from a query result set.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The fetched rows.

Raises:

ProgrammingError – If called before executing a query that returns results.

async fetchall() → list[Any | dict[Any, Any | None]][source]

Fetch all remaining rows from a query result set.

Returns:

The remaining rows.

Raises:

ProgrammingError – If called before executing a query that returns results.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

async cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

class pyathena.aio.cursor.AioDictCursor(dict_type: type[Any] | None = None, **kwargs)[source]

Native asyncio cursor that returns rows as dictionaries.

Example

>>> async with await AioConnection.create(...) as conn:
...     cursor = conn.cursor(AioDictCursor)
...     await cursor.execute("SELECT id, name FROM users")
...     row = await cursor.fetchone()
...     print(row["name"])
__init__(dict_type: type[Any] | None = None, **kwargs) → None[source]

Initialize an AioDictCursor.

Parameters:
  • dict_type – The type used to build each row of this cursor’s result sets. If None, the result set class’s dict_type is used.

  • **kwargs – Arguments forwarded to AioCursor.__init__.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

async cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → AioCursor

Execute a SQL query asynchronously.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters (optional).

  • work_group – Athena workgroup to use (optional).

  • s3_staging_dir – S3 location for query results (optional).

  • cache_size – Number of queries to check for result caching (optional).

  • cache_expiration_time – Cache expiration time in seconds (optional).

  • result_reuse_enable – Enable result reuse (optional).

  • result_reuse_minutes – Result reuse duration in minutes (optional).

  • paramstyle – Parameter style to use (optional).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

Self reference for method chaining.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

async fetchall() → list[Any | dict[Any, Any | None]]

Fetch all remaining rows from a query result set.

Returns:

The remaining rows.

Raises:

ProgrammingError – If called before executing a query that returns results.

async fetchmany(size: int | None = None) → list[Any | dict[Any, Any | None]]

Fetch multiple rows from a query result set.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The fetched rows.

Raises:

ProgrammingError – If called before executing a query that returns results.

async fetchone() → Any | dict[Any, Any | None] | None

Fetch the next row of a query result set.

Returns:

The next row (a tuple, or a dict for AioDictCursor), or None if no more rows.

Raises:

ProgrammingError – If called before executing a query that returns results.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio Result Set

class pyathena.aio.result_set.AthenaAioResultSet(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None)[source]

Async result set that provides async fetch methods.

Skips the synchronous _pre_fetch by passing _pre_fetch=False to the parent __init__ and provides an async create() classmethod factory instead. Synchronous iteration raises TypeError; use async for instead.

__init__(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None) → None[source]

Initialize the result set without fetching rows; create() fetches the first page.

Parameters:
  • connection – The connection that ran the query.

  • converter – The converter for result values.

  • query_execution – The query execution whose results to read.

  • arraysize – The number of rows per GetQueryResults page and the default fetchmany() size.

  • retry_config – The retry configuration for API calls.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

Raises:

ProgrammingError – If query_execution is not given.

async classmethod create(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None, **kwargs: Any) → AthenaAioResultSet[source]

Async factory method.

Creates an AthenaAioResultSet and awaits the initial data fetch.

Parameters:
  • connection – The database connection.

  • converter – Type converter for result values.

  • query_execution – Query execution metadata.

  • arraysize – Number of rows to fetch per request.

  • retry_config – Retry configuration for API calls.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • **kwargs – Additional arguments passed to the constructor of cls, such as dict_type for AthenaAioDictResultSet.

Returns:

A fully initialized AthenaAioResultSet.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None[source]

Fetch the next row of the result set.

Automatically fetches the next page from Athena when the current page is exhausted and more pages are available.

Returns:

The next row (a tuple, or a dict for AthenaAioDictResultSet), or None if no more rows.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]][source]

Fetch multiple rows from the result set.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The rows, fewer than size when the result is exhausted.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]][source]

Fetch all remaining rows from the result set.

Returns:

The remaining rows.

__iter__() → NoReturn[source]

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
property arraysize: int

The default number of rows per fetchmany() call.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the result set and discard its query execution, metadata, and rows.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection of the result set; raises ProgrammingError if closed.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions.

None without result metadata, or for INSERT, UPDATE, DELETE, and MERGE.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

property is_closed: bool

Whether the result set is closed.

property is_unload: bool

Check if the query is an UNLOAD statement.

Returns:

True if the query is an UNLOAD statement, False otherwise.

property kms_key: str | None

The KMS key used to encrypt the query results.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The ID of the query execution.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

The number of rows affected by the last operation, or -1 if it is unknown.

property rownumber: int | None

The zero-based index of the next row, or None if it is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

class pyathena.aio.result_set.AthenaAioDictResultSet(*args: Any, dict_type: type[Any] | None = None, **kwargs: Any)[source]

Async result set that returns rows as dictionaries.

Inherits _get_rows from AthenaDictResultSet and async fetch methods from AthenaAioResultSet via multiple inheritance.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
__init__(*args: Any, dict_type: type[Any] | None = None, **kwargs: Any) → None

Initialize the result set with an optional row type for this instance.

Parameters:
  • *args – Positional arguments passed to the next __init__ in the MRO.

  • dict_type – The type used to build each row of this result set. If None, the class attribute dict_type is used.

  • **kwargs – Keyword arguments passed to the next __init__ in the MRO.

__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

property arraysize: int

The default number of rows per fetchmany() call.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the result set and discard its query execution, metadata, and rows.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection of the result set; raises ProgrammingError if closed.

async classmethod create(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, result_set_type_hints: dict[str | int, str] | None = None, **kwargs: Any) → AthenaAioResultSet

Async factory method.

Creates an AthenaAioResultSet and awaits the initial data fetch.

Parameters:
  • connection – The database connection.

  • converter – Type converter for result values.

  • query_execution – Query execution metadata.

  • arraysize – Number of rows to fetch per request.

  • retry_config – Retry configuration for API calls.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • **kwargs – Additional arguments passed to the constructor of cls, such as dict_type for AthenaAioDictResultSet.

Returns:

A fully initialized AthenaAioResultSet.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions.

None without result metadata, or for INSERT, UPDATE, DELETE, and MERGE.

dict_type

alias of dict

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Returns:

The remaining rows.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The rows, fewer than size when the result is exhausted.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Automatically fetches the next page from Athena when the current page is exhausted and more pages are available.

Returns:

The next row (a tuple, or a dict for AthenaAioDictResultSet), or None if no more rows.

property is_closed: bool

Whether the result set is closed.

property is_unload: bool

Check if the query is an UNLOAD statement.

Returns:

True if the query is an UNLOAD statement, False otherwise.

property kms_key: str | None

The KMS key used to encrypt the query results.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The ID of the query execution.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

The number of rows affected by the last operation, or -1 if it is unknown.

property rownumber: int | None

The zero-based index of the next row, or None if it is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio Base Classes

class pyathena.aio.common.AioBaseCursor(connection: Connection[Any], converter: Converter, formatter: Formatter, retry_config: RetryConfig, s3_staging_dir: str | None, schema_name: str | None, catalog_name: str | None, work_group: str | None, poll_interval: float, encryption_option: str | None, kms_key: str | None, kill_on_interrupt: bool, result_reuse_enable: bool, result_reuse_minutes: int, on_start_query_execution: Callable[[str], None] | None = None, on_poll: OnPollCallback | None = None, **kwargs)[source]

Async base cursor that overrides I/O methods with async equivalents.

Reuses BaseCursor.__init__, all _build_* methods, and constants. Only the methods that perform network I/O or blocking sleep are overridden to use asyncio.to_thread / asyncio.sleep.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase][source]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata[source]

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata][source]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__init__(connection: Connection[Any], converter: Converter, formatter: Formatter, retry_config: RetryConfig, s3_staging_dir: str | None, schema_name: str | None, catalog_name: str | None, work_group: str | None, poll_interval: float, encryption_option: str | None, kms_key: str | None, kill_on_interrupt: bool, result_reuse_enable: bool, result_reuse_minutes: int, on_start_query_execution: Callable[[str], None] | None = None, on_poll: OnPollCallback | None = None, **kwargs) → None

Initialize the cursor with the settings it uses to run queries.

Parameters:
  • connection – The connection that created the cursor.

  • converter – Converter for result values.

  • formatter – Formatter for query parameters.

  • retry_config – Retry configuration for API calls.

  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).

  • kms_key – KMS key for encryption.

  • kill_on_interrupt – Cancel the execution when a KeyboardInterrupt interrupts starting it or waiting for it.

  • result_reuse_enable – Enable Athena query result reuse.

  • result_reuse_minutes – Maximum age in minutes of a reused result.

  • on_start_query_execution – Callback invoked with each query ID before the cursor waits for the query, by cursors whose execute() supports it.

  • on_poll – Callback invoked once per poll iteration with the current execution object.

  • **kwargs – Ignored.

abstractmethod close() → None

Close the cursor.

property connection: Connection[Any]

The connection that created this cursor.

abstractmethod execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, **kwargs)

Execute a SQL query.

Parameters:
  • operation – SQL query string.

  • parameters – Query parameters.

  • **kwargs – Execution options defined by the cursor implementation.

abstractmethod executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query once for each set of parameters.

Parameters:
  • operation – SQL query string.

  • seq_of_parameters – Sequence of parameter sets.

  • **kwargs – Execution options defined by the cursor implementation.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

class pyathena.aio.common.WithAsyncFetch(arraysize: int | None = None, **kwargs)[source]

Base class of the asyncio SQL cursors.

Combines WithResultSet with AioBaseCursor and CursorIterator, and provides async fetch, executemany, and cancel, async iteration, and the async context manager protocol. The fetch methods run the result set’s synchronous fetch with asyncio.to_thread; a subclass whose result set fetches asynchronously overrides them. Synchronous iteration raises TypeError.

Subclasses override execute() and optionally __init__ and format-specific helpers.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None[source]

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

async cancel() → None[source]

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None[source]

Fetch the next row of the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

A tuple representing the next row, or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]][source]

Fetch multiple rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

List of tuples representing the fetched rows.

Raises:

ProgrammingError – If no result set is available.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]][source]

Fetch all remaining rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

List of tuples representing all remaining rows.

Raises:

ProgrammingError – If no result set is available.

__iter__() → NoReturn[source]

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__init__(arraysize: int | None = None, **kwargs) → None

Initialize the cursor with no query ID and no result set.

Parameters:
  • arraysize – Default number of rows per fetchmany() call, validated by the arraysize setter. If None, DEFAULT_FETCH_SIZE is used.

  • **kwargs – Arguments passed to the next __init__ in the MRO.

Raises:

ProgrammingError – If arraysize is outside the range the cursor’s arraysize setter accepts.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

abstractmethod execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, **kwargs)

Execute a SQL query.

Parameters:
  • operation – SQL query string.

  • parameters – Query parameters.

  • **kwargs – Execution options defined by the cursor implementation.

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio Pandas Cursor

class pyathena.aio.pandas.cursor.AioPandasCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, engine: str = 'auto', chunksize: int | None = None, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, auto_optimize_chunksize: bool = False, **kwargs)[source]

Native asyncio cursor that returns results as pandas DataFrames.

Uses asyncio.to_thread() for both result set creation and fetch operations, keeping the event loop free. This is especially important when chunksize is set, as fetch calls trigger lazy S3 reads.

Example

>>> async with await pyathena.aio_connect(...) as conn:
...     cursor = conn.cursor(AioPandasCursor)
...     await cursor.execute("SELECT * FROM my_table")
...     df = cursor.as_pandas()
__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, engine: str = 'auto', chunksize: int | None = None, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, auto_optimize_chunksize: bool = False, **kwargs) → None[source]

Initialize an AioPandasCursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option for query results.

  • kms_key – KMS key for encrypting query results.

  • kill_on_interrupt – Cancel the query when the task is cancelled while execute() starts or waits for the query.

  • unload – Whether to wrap queries in UNLOAD and read the Parquet output.

  • engine – Parsing engine (auto, c, python, or pyarrow).

  • chunksize – Number of rows per DataFrame chunk when reading CSV results. If set, it takes precedence over auto_optimize_chunksize.

  • block_size – Default block size of the S3 filesystem that reads the results.

  • cache_type – Default cache type of the S3 filesystem that reads the results.

  • max_workers – Maximum number of workers of the S3 filesystem.

  • result_reuse_enable – Whether to enable Athena query result reuse.

  • result_reuse_minutes – Maximum age of a reused query result in minutes.

  • auto_optimize_chunksize – Whether to choose a chunk size from the size of the CSV result file when chunksize is None.

  • **kwargs – Other cursor arguments, such as connection and arraysize, passed to the parent __init__.

static get_default_converter(unload: bool = False) → DefaultPandasTypeConverter | Any[source]

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, keep_default_na: bool = False, na_values: Iterable[str] | None = ('',), quoting: int = 1, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → AioPandasCursor[source]

Execute a SQL query asynchronously and return results as pandas DataFrames.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters for parameterized queries.

  • work_group – Athena workgroup to use for this query.

  • s3_staging_dir – S3 location for query results.

  • cache_size – Number of queries to check for result caching.

  • cache_expiration_time – Cache expiration time in seconds.

  • result_reuse_enable – Enable Athena result reuse for this query.

  • result_reuse_minutes – Minutes to reuse cached results.

  • paramstyle – Parameter style (‘qmark’ or ‘pyformat’).

  • keep_default_na – Whether to keep default pandas NA values.

  • na_values – Additional values to treat as NA.

  • quoting – CSV quoting behavior (pandas csv.QUOTE_* constants).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional pandas read_csv/read_parquet parameters.

Returns:

Self reference for method chaining.

as_pandas() → DataFrame | PandasDataFrameIterator[source]

Return DataFrame or PandasDataFrameIterator based on chunksize setting.

Returns:

DataFrame when chunksize is None, PandasDataFrameIterator when chunksize is set.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

async cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

List of tuples representing all remaining rows.

Raises:

ProgrammingError – If no result set is available.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

List of tuples representing the fetched rows.

Raises:

ProgrammingError – If no result set is available.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

A tuple representing the next row, or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio Arrow Cursor

class pyathena.aio.arrow.cursor.AioArrowCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, connect_timeout: float | None = None, request_timeout: float | None = None, **kwargs)[source]

Native asyncio cursor that returns results as Apache Arrow Tables.

Uses asyncio.to_thread() for both result set creation and fetch operations, keeping the event loop free.

Example

>>> async with await pyathena.aio_connect(...) as conn:
...     cursor = conn.cursor(AioArrowCursor)
...     await cursor.execute("SELECT * FROM my_table")
...     table = cursor.as_arrow()
__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, connect_timeout: float | None = None, request_timeout: float | None = None, **kwargs) → None[source]

Initialize an AioArrowCursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option for query results.

  • kms_key – KMS key for encrypting query results.

  • kill_on_interrupt – Cancel the query when the task is cancelled while execute() starts or waits for the query.

  • unload – Whether to wrap queries in UNLOAD and read the Parquet output.

  • result_reuse_enable – Whether to enable Athena query result reuse.

  • result_reuse_minutes – Maximum age of a reused query result in minutes.

  • connect_timeout – Connection timeout in seconds of the pyarrow S3 filesystem that reads the results. If None, the pyarrow default is used.

  • request_timeout – Request timeout in seconds of the pyarrow S3 filesystem that reads the results. If None, the pyarrow default is used.

  • **kwargs – Other cursor arguments, such as connection and arraysize, passed to the parent __init__.

static get_default_converter(unload: bool = False) → DefaultArrowTypeConverter | DefaultArrowUnloadTypeConverter | Any[source]

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → AioArrowCursor[source]

Execute a SQL query asynchronously and return results as Arrow Tables.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters for parameterized queries.

  • work_group – Athena workgroup to use for this query.

  • s3_staging_dir – S3 location for query results.

  • cache_size – Number of queries to check for result caching.

  • cache_expiration_time – Cache expiration time in seconds.

  • result_reuse_enable – Enable Athena result reuse for this query.

  • result_reuse_minutes – Minutes to reuse cached results.

  • paramstyle – Parameter style (‘qmark’ or ‘pyformat’).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

Self reference for method chaining.

as_arrow() → Table[source]

Return query results as an Apache Arrow Table.

Returns:

Apache Arrow Table containing all query results.

as_polars() → pl.DataFrame[source]

Return query results as a Polars DataFrame.

Returns:

Polars DataFrame containing all query results.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

async cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

List of tuples representing all remaining rows.

Raises:

ProgrammingError – If no result set is available.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

List of tuples representing the fetched rows.

Raises:

ProgrammingError – If no result set is available.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

A tuple representing the next row, or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio Polars Cursor

class pyathena.aio.polars.cursor.AioPolarsCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, chunksize: int | None = None, **kwargs)[source]

Native asyncio cursor that returns results as Polars DataFrames.

Uses asyncio.to_thread() for both result set creation and fetch operations, keeping the event loop free. This is especially important when chunksize is set, as fetch calls trigger lazy S3 reads.

Example

>>> async with await pyathena.aio_connect(...) as conn:
...     cursor = conn.cursor(AioPolarsCursor)
...     await cursor.execute("SELECT * FROM my_table")
...     df = cursor.as_polars()
__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, unload: bool = False, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, block_size: int | None = None, cache_type: str | None = None, max_workers: int = 20, chunksize: int | None = None, **kwargs) → None[source]

Initialize an AioPolarsCursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option for query results.

  • kms_key – KMS key for encrypting query results.

  • kill_on_interrupt – Cancel the query when the task is cancelled while execute() starts or waits for the query.

  • unload – Whether to wrap queries in UNLOAD and read the Parquet output.

  • result_reuse_enable – Whether to enable Athena query result reuse.

  • result_reuse_minutes – Maximum age of a reused query result in minutes.

  • block_size – Default block size of the S3 filesystem that reads the results.

  • cache_type – Default cache type of the S3 filesystem that reads the results.

  • max_workers – Maximum number of workers of the S3 filesystem.

  • chunksize – Number of rows per chunk. If set, result files in S3 are read lazily in chunks of this size.

  • **kwargs – Other cursor arguments, such as connection and arraysize, passed to the parent __init__.

static get_default_converter(unload: bool = False) → DefaultPolarsTypeConverter | DefaultPolarsUnloadTypeConverter | Any[source]

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → AioPolarsCursor[source]

Execute a SQL query asynchronously and return results as Polars DataFrames.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters for parameterized queries.

  • work_group – Athena workgroup to use for this query.

  • s3_staging_dir – S3 location for query results.

  • cache_size – Number of queries to check for result caching.

  • cache_expiration_time – Cache expiration time in seconds.

  • result_reuse_enable – Enable Athena result reuse for this query.

  • result_reuse_minutes – Minutes to reuse cached results.

  • paramstyle – Parameter style (‘qmark’ or ‘pyformat’).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters passed to Polars read functions.

Returns:

Self reference for method chaining.

as_polars() → pl.DataFrame[source]

Return query results as a Polars DataFrame.

Returns:

Polars DataFrame containing all query results.

as_arrow() → Table[source]

Return query results as an Apache Arrow Table.

Returns:

Apache Arrow Table containing all query results.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

async cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

List of tuples representing all remaining rows.

Raises:

ProgrammingError – If no result set is available.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

List of tuples representing the fetched rows.

Raises:

ProgrammingError – If no result set is available.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

A tuple representing the next row, or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio S3FS Cursor

class pyathena.aio.s3fs.cursor.AioS3FSCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, csv_reader: type[DefaultCSVReader] | type[AthenaCSVReader] | None = None, **kwargs)[source]

Native asyncio cursor that reads CSV results via AioS3FileSystem.

Uses AioS3FileSystem for S3 operations, which replaces ThreadPoolExecutor parallelism with asyncio.gather + asyncio.to_thread. Fetch operations are wrapped in asyncio.to_thread() because CSV reading is blocking I/O.

Example

>>> async with await pyathena.aio_connect(...) as conn:
...     cursor = conn.cursor(AioS3FSCursor)
...     await cursor.execute("SELECT * FROM my_table")
...     row = await cursor.fetchone()
__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, csv_reader: type[DefaultCSVReader] | type[AthenaCSVReader] | None = None, **kwargs) → None[source]

Initialize an AioS3FSCursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option for query results.

  • kms_key – KMS key for encrypting query results.

  • kill_on_interrupt – Cancel the query when the task is cancelled while execute() starts or waits for the query.

  • result_reuse_enable – Whether to enable Athena query result reuse.

  • result_reuse_minutes – Maximum age of a reused query result in minutes.

  • csv_reader – CSV reader class for parsing the result files. If None, AthenaCSVReader is used, which distinguishes NULL from empty strings. DefaultCSVReader reads both as empty strings.

  • **kwargs – Other cursor arguments, such as connection and arraysize, passed to the parent __init__.

static get_default_converter(unload: bool = False) → DefaultS3FSTypeConverter[source]

Get the default type converter for S3FS cursor.

Parameters:

unload – Unused. S3FS cursor does not support UNLOAD operations.

Returns:

DefaultS3FSTypeConverter instance.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → AioS3FSCursor[source]

Execute a SQL query asynchronously via S3FileSystem CSV reader.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters for parameterized queries.

  • work_group – Athena workgroup to use for this query.

  • s3_staging_dir – S3 location for query results.

  • cache_size – Number of queries to check for result caching.

  • cache_expiration_time – Cache expiration time in seconds.

  • result_reuse_enable – Enable Athena result reuse for this query.

  • result_reuse_minutes – Minutes to reuse cached results.

  • paramstyle – Parameter style (‘qmark’ or ‘pyformat’).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

Self reference for method chaining.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__iter__() → NoReturn

Reject synchronous iteration; use async for instead.

Raises:

TypeError – Always, because the fetch methods are coroutines.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

async cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure or cancellation, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

async fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

List of tuples representing all remaining rows.

Raises:

ProgrammingError – If no result set is available.

async fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

List of tuples representing the fetched rows.

Raises:

ProgrammingError – If no result set is available.

async fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Wraps the synchronous fetch in asyncio.to_thread to avoid blocking the event loop.

Returns:

A tuple representing the next row, or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

async get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

async list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

async list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Aio Spark Cursor

class pyathena.aio.spark.cursor.AioSparkCursor(session_id: str | None = None, description: str | None = None, engine_configuration: dict[str, Any] | None = None, notebook_version: str | None = None, session_idle_timeout_minutes: int | None = None, terminate_session_on_close: bool | None = None, **kwargs)[source]

Native asyncio cursor for executing PySpark code on Athena.

Overrides post-init I/O methods of SparkBaseCursor with async equivalents. Session management (_exists_session, _start_session, etc.) stays synchronous because __init__ runs inside asyncio.to_thread.

Since SparkBaseCursor.__init__ performs I/O (session management), cursor creation must be wrapped in asyncio.to_thread:

cursor = await asyncio.to_thread(conn.cursor)

Example

>>> import asyncio
>>> async with await pyathena.aio_connect(
...     work_group="spark-workgroup",
...     cursor_class=AioSparkCursor,
... ) as conn:
...     cursor = await asyncio.to_thread(conn.cursor)
...     await cursor.execute("spark.sql('SELECT 1').show()")
...     print(await cursor.get_std_out())
property calculation_execution: AthenaCalculationExecution | None

The calculation execution that the other properties read, or None.

async get_std_out() → str | None[source]

Get the standard output from the Spark calculation execution.

Returns:

The standard output as a string, or None if no output is available.

async get_std_error() → str | None[source]

Get the standard error from the Spark calculation execution.

Returns:

The standard error as a string, or None if no error output is available.

async execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, session_id: str | None = None, description: str | None = None, client_request_token: str | None = None, work_group: str | None = None, **kwargs) → AioSparkCursor[source]

Execute PySpark code asynchronously.

Parameters:
  • operation – PySpark code to execute.

  • parameters – Unused, kept for API compatibility.

  • session_id – Spark session ID override.

  • description – Calculation description.

  • client_request_token – Idempotency token.

  • work_group – Unused, kept for API compatibility.

  • **kwargs – Additional parameters.

Returns:

Self reference for method chaining.

async cancel() → None[source]

Cancel the currently running calculation.

Raises:

ProgrammingError – If no calculation is running.

async close() → None[source]

Close the cursor, terminating its Spark session if configured to.

See terminate_session_on_close. After a successful termination, further calls do not terminate the session again; after a failed one, calling this method again retries it.

Raises:

OperationalError – If terminating the session fails.

async executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None[source]

Execute a SQL query once for each set of parameters.

Parameters:
  • operation – SQL query string.

  • seq_of_parameters – Sequence of parameter sets.

  • **kwargs – Execution options defined by the cursor implementation.

LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
__init__(session_id: str | None = None, description: str | None = None, engine_configuration: dict[str, Any] | None = None, notebook_version: str | None = None, session_idle_timeout_minutes: int | None = None, terminate_session_on_close: bool | None = None, **kwargs) → None

Initialize the cursor and start or attach to a Spark session.

If waiting for a newly started session fails, that session is terminated regardless of terminate_session_on_close; a supplied session is not.

Parameters:
  • session_id – ID of an existing session to use. If omitted, a new session is started.

  • description – Description of a new session.

  • engine_configuration – Engine configuration of a new session. Defaults to get_default_engine_configuration().

  • notebook_version – Notebook version of a new session.

  • session_idle_timeout_minutes – Idle timeout of a new session in minutes.

  • terminate_session_on_close – Whether close() terminates the session. If None, only a session started by this cursor is terminated; a session supplied with session_id is left running.

  • **kwargs – Arguments passed to BaseCursor.

Raises:

OperationalError – If the supplied session does not exist, or the session cannot be started or does not become idle.

property calculation_id: str | None

The ID of the calculation tracked by this cursor, or None if there is none.

property completion_date_time: datetime | None

The CompletionDateTime of the calculation, or None if there is none.

property connection: Connection[Any]

The connection that created this cursor.

property description: str | None

The Description of the calculation, or None if there is none.

property dpu_execution_in_millis: int | None

The DpuExecutionInMillis statistic of the calculation, or None if there is none.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

static get_default_engine_configuration() → dict[str, Any]

Return the engine configuration used when none is given.

Returns:

a coordinator DPU size of 1, at most 2 concurrent DPUs, and a default executor DPU size of 1.

Return type:

The EngineConfiguration of a new session

get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property progress: str | None

The Progress statistic of the calculation, or None if there is none.

property result_s3_uri: str | None

The ResultS3Uri of the calculation, or None if there is none.

property result_type: str | None

The ResultType of the calculation, or None if there is none.

property session_id: str

The ID of the Spark session that this cursor runs calculations in.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The State of the calculation, or None if there is none.

property state_change_reason: str | None

The StateChangeReason of the calculation, or None if there is none.

property std_error_s3_uri: str | None

The StdErrorS3Uri of the calculation, or None if there is none.

property std_out_s3_uri: str | None

The StdOutS3Uri of the calculation, or None if there is none.

property submission_date_time: datetime | None

The SubmissionDateTime of the calculation, or None if there is none.

property working_directory: str | None

The WorkingDirectory of the calculation, or None if there is none.