Connection and Cursors

This section covers the core connection functionality and basic cursor operations.

Connection

DB API 2.0 interface to Amazon Athena: connect(), aio_connect(), and type objects.

class pyathena.DBAPITypeObject[source]

A DB API type object that compares equal to each of its Athena type names.

https://www.python.org/dev/peps/pep-0249/#type-objects-and-constructors

pyathena.connect(*args, cursor_class: None = ..., **kwargs) → Connection[Cursor][source]
pyathena.connect(*args, cursor_class: type[ConnectionCursor], **kwargs) → Connection[ConnectionCursor]

Create a new database connection to Amazon Athena.

This function provides the main entry point for establishing connections to Amazon Athena. It follows the DB API 2.0 specification and returns a Connection object that can be used to create cursors for executing SQL queries.

Parameters:
  • *args – Positional arguments passed to the Connection constructor, in the order of its parameters (s3_staging_dir, region_name, …).

  • s3_staging_dir – S3 location to store query results. Required if not using workgroups or if the workgroup doesn’t have a result location. Pass an empty string to explicitly disable S3 staging and skip the AWS_ATHENA_S3_STAGING_DIR environment variable fallback (required for workgroups with managed query result storage).

  • region_name – AWS region name. If not specified, uses the default region from your AWS configuration.

  • schema_name – Athena database/schema name. Defaults to “default”.

  • catalog_name – Athena data catalog name. Defaults to “awsdatacatalog”.

  • work_group – Athena workgroup name. Can be used instead of s3_staging_dir if the workgroup has a result location configured.

  • poll_interval – Time in seconds between polling for query completion. Defaults to 1.0.

  • encryption_option – S3 encryption option for query results. Can be “SSE_S3”, “SSE_KMS”, or “CSE_KMS”.

  • kms_key – KMS key ID for encryption when using SSE_KMS or CSE_KMS.

  • profile_name – AWS profile name to use for authentication.

  • role_arn – ARN of IAM role to assume for authentication.

  • role_session_name – Session name when assuming a role.

  • cursor_class – Custom cursor class to use. If not specified, uses the default Cursor class.

  • kill_on_interrupt – Whether to cancel running queries when interrupted. Defaults to True.

  • **kwargs – Additional keyword arguments passed to the Connection constructor.

Returns:

A Connection object that can be used to create cursors and execute queries.

Raises:

ProgrammingError – If neither s3_staging_dir nor work_group is provided.

Example

>>> import pyathena
>>> conn = pyathena.connect(
...     s3_staging_dir='s3://my-bucket/staging/',
...     region_name='us-east-1',
...     schema_name='mydatabase'
... )
>>> cursor = conn.cursor()
>>> cursor.execute("SELECT * FROM mytable LIMIT 10")
>>> results = cursor.fetchall()
async pyathena.aio_connect(*args, **kwargs) → AioConnection[source]

Create a new async database connection to Amazon Athena.

This is the async counterpart of connect(). It returns an AioConnection whose cursors use native asyncio for polling and API calls, keeping the event loop free.

Parameters:
  • *args – Forwarded to AioConnection.create(), which accepts keyword arguments only.

  • **kwargs – Arguments forwarded to AioConnection.create(). See connect() for the full list of supported arguments.

Returns:

An AioConnection that produces AioCursor instances by default.

Example

>>> import pyathena
>>> conn = await pyathena.aio_connect(
...     s3_staging_dir='s3://my-bucket/staging/',
...     region_name='us-east-1',
... )
>>> async with conn.cursor() as cursor:
...     await cursor.execute("SELECT 1")
...     print(await cursor.fetchone())
class pyathena.connection.Connection(s3_staging_dir: str | None = ..., region_name: str | None = ..., schema_name: str | None = ..., catalog_name: str | None = ..., work_group: str | None = ..., poll_interval: float = ..., encryption_option: str | None = ..., kms_key: str | None = ..., profile_name: str | None = ..., role_arn: str | None = ..., role_session_name: str = ..., external_id: str | None = ..., serial_number: str | None = ..., duration_seconds: int = ..., converter: Converter | None = ..., formatter: Formatter | None = ..., retry_config: RetryConfig | None = ..., cursor_class: None = ..., cursor_kwargs: dict[str, Any] | None = ..., kill_on_interrupt: bool = ..., session: Session | None = ..., config: Config | None = ..., result_reuse_enable: bool = ..., result_reuse_minutes: int = ..., on_start_query_execution: Callable[[str], None] | None = ..., on_poll: Callable[[AthenaQueryExecution | AthenaCalculationExecutionStatus], None] | None = ..., glue_metadata_fallback: bool = ..., **kwargs)[source]
class pyathena.connection.Connection(s3_staging_dir: str | None = ..., region_name: str | None = ..., schema_name: str | None = ..., catalog_name: str | None = ..., work_group: str | None = ..., poll_interval: float = ..., encryption_option: str | None = ..., kms_key: str | None = ..., profile_name: str | None = ..., role_arn: str | None = ..., role_session_name: str = ..., external_id: str | None = ..., serial_number: str | None = ..., duration_seconds: int = ..., converter: Converter | None = ..., formatter: Formatter | None = ..., retry_config: RetryConfig | None = ..., cursor_class: type[ConnectionCursor] = ..., cursor_kwargs: dict[str, Any] | None = ..., kill_on_interrupt: bool = ..., session: Session | None = ..., config: Config | None = ..., result_reuse_enable: bool = ..., result_reuse_minutes: int = ..., on_start_query_execution: Callable[[str], None] | None = ..., on_poll: Callable[[AthenaQueryExecution | AthenaCalculationExecutionStatus], None] | None = ..., glue_metadata_fallback: bool = ..., **kwargs)

A DB API 2.0 compliant connection to Amazon Athena.

The Connection class represents a database session and provides methods to create cursors for executing SQL queries against Amazon Athena. It handles authentication, session management, and query result storage in S3.

This class follows the Python Database API Specification v2.0 (PEP 249) and provides a familiar interface for database operations.

s3_staging_dir

S3 location where query results are stored.

region_name

AWS region name.

schema_name

Default database/schema name for queries.

catalog_name

Data catalog name (typically “awsdatacatalog”).

work_group

Athena workgroup name.

poll_interval

Interval in seconds for polling query status.

encryption_option

S3 encryption option for query results.

kms_key

KMS key for encryption when applicable.

kill_on_interrupt

Whether to cancel queries on interrupt signals.

result_reuse_enable

Whether to enable Athena’s result reuse feature.

result_reuse_minutes

Minutes to reuse cached results.

Example

>>> conn = Connection(
...     s3_staging_dir='s3://my-bucket/staging/',
...     region_name='us-east-1',
...     schema_name='mydatabase'
... )
>>> with conn:
...     cursor = conn.cursor()
...     cursor.execute("SELECT COUNT(*) FROM mytable")
...     result = cursor.fetchone()

Note

Either s3_staging_dir or work_group must be specified. If using a workgroup, it must have a result location configured unless s3_staging_dir is also provided. For workgroups with managed query result storage, pass s3_staging_dir="" to skip the environment variable fallback.

__init__(s3_staging_dir: str | None = None, region_name: str | None = None, schema_name: str | None = 'default', catalog_name: str | None = 'awsdatacatalog', work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, profile_name: str | None = None, role_arn: str | None = None, role_session_name: str = 'PyAthena-session-1791030925', external_id: str | None = None, serial_number: str | None = None, duration_seconds: int = 3600, converter: Converter | None = None, formatter: Formatter | None = None, retry_config: RetryConfig | None = None, cursor_class: None = None, cursor_kwargs: dict[str, Any] | None = None, kill_on_interrupt: bool = True, session: Session | None = None, config: Config | None = None, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, on_start_query_execution: Callable[[str], None] | None = None, on_poll: Callable[[AthenaQueryExecution | AthenaCalculationExecutionStatus], None] | None = None, glue_metadata_fallback: bool = True, **kwargs) → None[source]
__init__(s3_staging_dir: str | None = None, region_name: str | None = None, schema_name: str | None = 'default', catalog_name: str | None = 'awsdatacatalog', work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, profile_name: str | None = None, role_arn: str | None = None, role_session_name: str = 'PyAthena-session-1791030925', external_id: str | None = None, serial_number: str | None = None, duration_seconds: int = 3600, converter: Converter | None = None, formatter: Formatter | None = None, retry_config: RetryConfig | None = None, cursor_class: type[ConnectionCursor] = None, cursor_kwargs: dict[str, Any] | None = None, kill_on_interrupt: bool = True, session: Session | None = None, config: Config | None = None, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, on_start_query_execution: Callable[[str], None] | None = None, on_poll: Callable[[AthenaQueryExecution | AthenaCalculationExecutionStatus], None] | None = None, glue_metadata_fallback: bool = True, **kwargs) → None

Initialize a new Athena database connection.

Parameters:
  • s3_staging_dir – S3 location to store query results. Required if not using workgroups or if workgroup doesn’t have result location. Pass an empty string to explicitly disable S3 staging and skip the AWS_ATHENA_S3_STAGING_DIR environment variable fallback. This is required when connecting to a workgroup with managed query result storage enabled.

  • region_name – AWS region name. Uses default region if not specified.

  • schema_name – Default database/schema name. Defaults to “default”.

  • catalog_name – Data catalog name. Defaults to “awsdatacatalog”.

  • work_group – Athena workgroup name. Can substitute for s3_staging_dir if workgroup has result location configured.

  • poll_interval – Seconds between query status polls. Defaults to 1.0.

  • encryption_option – S3 encryption for results (“SSE_S3”, “SSE_KMS”, “CSE_KMS”).

  • kms_key – KMS key ID when using SSE_KMS or CSE_KMS encryption.

  • profile_name – AWS profile name for authentication.

  • role_arn – IAM role ARN to assume for authentication.

  • role_session_name – Session name when assuming IAM role.

  • external_id – External ID for role assumption (if required by role).

  • serial_number – MFA device serial number for role assumption.

  • duration_seconds – Role session duration in seconds. Defaults to 3600.

  • converter – Custom type converter. Uses DefaultTypeConverter if None.

  • formatter – Custom parameter formatter. Uses DefaultParameterFormatter if None.

  • retry_config – Retry configuration for API calls. Uses default if None.

  • cursor_class – Default cursor class for this connection.

  • cursor_kwargs – Default keyword arguments for cursor creation. Arguments passed to cursor() override the same keys.

  • kill_on_interrupt – Cancel running queries on interrupt. Defaults to True.

  • session – Pre-configured boto3 Session. Creates new session if None.

  • config – Boto3 Config object for client configuration.

  • result_reuse_enable – Enable Athena query result reuse. Defaults to False.

  • result_reuse_minutes – Minutes to reuse cached results.

  • on_start_query_execution – Callback invoked with each query ID before the cursor waits for the query, as for the execute() argument of the same name.

  • on_poll – Callback invoked once per poll iteration with the current execution object (AthenaQueryExecution, or AthenaCalculationExecutionStatus for Spark). Useful for monitoring live query progress. Defaults to None.

  • glue_metadata_fallback – In AwsDataCatalog and S3 Tables catalogs, answer a throttled table-metadata, table-listing or database-listing request from the AWS Glue Data Catalog before retrying it. Defaults to True.

  • **kwargs – Additional arguments passed to boto3 Session and client.

Raises:

ProgrammingError – If neither s3_staging_dir nor work_group is provided.

Note

Either s3_staging_dir or work_group must be specified. Environment variables AWS_ATHENA_S3_STAGING_DIR and AWS_ATHENA_WORK_GROUP are checked if parameters are not provided.

When using a workgroup with managed query result storage, pass s3_staging_dir="" to prevent the environment variable fallback from sending a ResultConfiguration that conflicts with ManagedQueryResultsConfiguration.

property session: Session

Get the boto3 session used for AWS API calls.

Returns:

The configured boto3 Session object.

property client: BaseClient

Get the boto3 Athena client used for query operations.

Returns:

The configured boto3 Athena client.

property retry_config: RetryConfig

Get the retry configuration for AWS API calls.

Returns:

The RetryConfig object that controls retry behavior for failed requests.

__enter__()[source]

Enter the runtime context for the connection.

Returns:

Self for use in context manager protocol.

__exit__(exc_type, exc_val, exc_tb)[source]

Exit the runtime context and close the connection.

Parameters:
  • exc_type – Exception type if an exception occurred.

  • exc_val – Exception value if an exception occurred.

  • exc_tb – Exception traceback if an exception occurred.

cursor(cursor: None = None, **kwargs) → ConnectionCursor[source]
cursor(cursor: type[FunctionalCursor], **kwargs) → FunctionalCursor

Create a new cursor object for executing queries.

Creates and returns a cursor object that can be used to execute SQL queries against Amazon Athena. The cursor inherits connection settings but can be customized with additional parameters.

Parameters:
  • cursor – Custom cursor class to use. If not provided, uses the connection’s default cursor class.

  • **kwargs – Additional keyword arguments to pass to the cursor constructor. These override the connection’s cursor_kwargs and its defaults.

Returns:

A cursor object that can execute SQL queries.

Example

>>> cursor = connection.cursor()
>>> cursor.execute("SELECT * FROM my_table LIMIT 10")
>>> results = cursor.fetchall()

# Using a custom cursor type >>> from pyathena.pandas.cursor import PandasCursor >>> pandas_cursor = connection.cursor(PandasCursor) >>> df = pandas_cursor.execute(“SELECT * FROM my_table”).fetchall()

close() → None[source]

Close the connection.

Closes the database connection. This method is provided for DB API 2.0 compatibility. Since Athena connections are stateless, this method currently does not perform any actual cleanup operations.

Note

This method is called automatically when using the connection as a context manager (with statement).

commit() → None[source]

Commit any pending transaction.

This method is provided for DB API 2.0 compatibility. Since Athena does not support transactions, this method does nothing.

Note

Athena queries are auto-committed and cannot be rolled back.

rollback() → None[source]

Rollback any pending transaction.

This method is required by DB API 2.0 but is not supported by Athena since Athena does not support transactions.

Raises:

NotSupportedError – Always raised since transactions are not supported.

Execution Options

class pyathena.options.ExecuteOptions(work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int = 0, cache_expiration_time: int = 0, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None)[source]

Shared options for Cursor.execute() across all cursor implementations.

This dataclass is the single source of truth for the query-execution arguments shared by every SQL cursor type (sync/async/aio and pandas/arrow/polars/s3fs variants). It can be passed to execute() via the options keyword argument as an alternative to individual :keyword >>> from pyathena.options import ExecuteOptions: :keyword >>> options = ExecuteOptions: :kwtype >>> options = ExecuteOptions: work_group=”primary”, cache_size=100 :keyword >>> cursor.execute: :kwtype >>> cursor.execute: “SELECT * FROM my_table”, options=options

When both options and individual keyword arguments are provided, the individual keyword arguments take precedence. This allows building a base ExecuteOptions once and tweaking it per call:

>>> cursor.execute("SELECT ...", options=options, work_group="adhoc")

Passing None for an individual keyword argument is treated as “not provided” and leaves the corresponding options field unchanged. merge() also ignores None, so to clear a field, use dataclasses.replace() or construct a new instance.

work_group

Athena workgroup to use for this query. Overrides the connection-level workgroup.

Type:

str | None

s3_staging_dir

S3 location for query results. Overrides the connection-level staging directory.

Type:

str | None

cache_size

Number of recent queries to scan for client-side result caching. 0 (default) disables the cache lookup, unless cache_expiration_time is set to a positive value, in which case all queries within the expiration window are scanned. A qmark query with parameters is never looked up.

Type:

int

cache_expiration_time

Maximum age in seconds of a cached query result to consider for reuse. 0 (default) means no age limit.

Type:

int

result_reuse_enable

Enable Athena server-side result reuse for this query. None (default) falls back to the connection-level setting.

Type:

bool | None

result_reuse_minutes

Maximum age in minutes of a previous query result that Athena should consider for reuse. None (default) falls back to the connection-level setting.

Type:

int | None

paramstyle

Parameter style for this query (‘qmark’ or ‘pyformat’). None (default) uses the module-level pyathena.paramstyle.

Type:

str | None

on_start_query_execution

Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution API call, or after a reusable query ID is found through cache_size. Invoked by synchronous and aio cursors; AsyncCursor-based cursors return the query ID directly through their execution model and do not invoke it.

Type:

collections.abc.Callable[[str], None] | None

result_set_type_hints

Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index. For example: {"tags": "array(varchar)", "metadata": "map(varchar, integer)"}

Type:

dict[str | int, str] | None

work_group: str | None = None
s3_staging_dir: str | None = None
cache_size: int = 0
cache_expiration_time: int = 0
result_reuse_enable: bool | None = None
result_reuse_minutes: int | None = None
paramstyle: str | None = None
on_start_query_execution: Callable[[str], None] | None = None
result_set_type_hints: dict[str | int, str] | None = None
classmethod resolve(options: ExecuteOptions | None, **overrides: Any) → ExecuteOptions[source]

Return options (or a default instance) with overrides applied.

This is the canonical way for execute() implementations to combine the options argument with the individual keyword arguments.

Parameters:
  • options – Base options, or None to start from the defaults.

  • **overrides – Field values to apply on top of options. None values are ignored.

Returns:

The effective ExecuteOptions for the call.

merge(**overrides: Any) → ExecuteOptions[source]

Return a new instance with non-None overrides applied.

Parameters:

**overrides – Field values to apply on top of this instance. None values are ignored, so an omitted execute() keyword argument never clobbers a value set on options.

Returns:

A new ExecuteOptions with the overrides applied.

Raises:

TypeError – If an override name is not a field of this class.

__init__(work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int = 0, cache_expiration_time: int = 0, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None) → None

Standard Cursors

class pyathena.cursor.Cursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs)[source]

A DB API 2.0 compliant cursor for executing SQL queries on Amazon Athena.

The Cursor class provides methods for executing SQL queries against Amazon Athena and retrieving results. It follows the Python Database API Specification v2.0 (PEP 249) and provides familiar database cursor operations.

This cursor returns results as tuples by default. For other data formats, consider using specialized cursor classes like PandasCursor or ArrowCursor.

description

Sequence of column descriptions for the last query.

rowcount

Number of rows affected by the last query (-1 for SELECT queries).

arraysize

Default number of rows to fetch with fetchmany().

Example

>>> cursor = connection.cursor()
>>> cursor.execute("SELECT name, age FROM users WHERE age > %(age)s", {"age": 18})
>>> while True:
...     row = cursor.fetchone()
...     if not row:
...         break
...     print(f"Name: {row[0]}, Age: {row[1]}")
>>> cursor.execute("CREATE TABLE test AS SELECT 1 as id, 'test' as name")
>>> print(f"Created table, rows affected: {cursor.rowcount}")
__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs) → None[source]

Initialize a Cursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).

  • kms_key – KMS key for encryption.

  • kill_on_interrupt – Cancel the query when a KeyboardInterrupt interrupts execute() while it starts or waits for the query.

  • result_reuse_enable – Enable Athena query result reuse.

  • result_reuse_minutes – Maximum age in minutes of a reused result.

  • **kwargs – Arguments forwarded to WithResultSet.__init__ and BaseCursor.__init__, such as arraysize, connection, converter, formatter, and retry_config.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → Cursor[source]

Execute a SQL query.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters (optional).

  • work_group – Athena workgroup to use for this query.

  • s3_staging_dir – S3 location for query results.

  • cache_size – Number of queries to check for result caching.

  • cache_expiration_time – Cache expiration time in seconds.

  • result_reuse_enable – Enable Athena result reuse for this query.

  • result_reuse_minutes – Minutes to reuse cached results.

  • paramstyle – Parameter style (‘qmark’ or ‘pyformat’).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size. Function signature: (query_id: str) -> None This allows early access to query_id for monitoring/cancellation.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index. For example: {"tags": "array(varchar)", "metadata": "map(varchar, integer)"}

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

Self reference for method chaining.

Example

>>> cursor.execute(
...     "SELECT * FROM table_with_complex_types",
...     result_set_type_hints={
...         "tags": "array(varchar)",
...         "metadata": "map(varchar, integer)",
...     }
... )
DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Returns:

The remaining rows.

Raises:

ProgrammingError – If no result set is available.

fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The fetched rows.

Raises:

ProgrammingError – If no result set is available.

fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Returns:

The next row (a tuple, or a dict for dict cursors), or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

class pyathena.cursor.DictCursor(dict_type: type[Any] | None = None, **kwargs)[source]

A cursor that returns query results as dictionaries instead of tuples.

DictCursor provides the same functionality as the standard Cursor but returns rows as dictionaries where column names are keys. This makes it easier to access column values by name rather than position.

Example

>>> cursor = connection.cursor(DictCursor)
>>> cursor.execute("SELECT id, name, email FROM users LIMIT 1")
>>> row = cursor.fetchone()
>>> print(f"User: {row['name']} ({row['email']})")
>>> cursor.execute("SELECT * FROM products")
>>> for row in cursor.fetchall():
...     print(f"Product {row['id']}: {row['name']} - ${row['price']}")
__init__(dict_type: type[Any] | None = None, **kwargs) → None[source]

Initialize a DictCursor.

Parameters:
  • dict_type – The type used to build each row of this cursor’s result sets. If None, the result set class’s dict_type is used.

  • **kwargs – Arguments forwarded to Cursor.__init__.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

cancel() → None

Cancel the currently executing query.

Raises:

ProgrammingError – If no query is currently executing.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the cursor and release associated resources.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection that created this cursor.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, on_start_query_execution: Callable[[str], None] | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → Cursor

Execute a SQL query.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters (optional).

  • work_group – Athena workgroup to use for this query.

  • s3_staging_dir – S3 location for query results.

  • cache_size – Number of queries to check for result caching.

  • cache_expiration_time – Cache expiration time in seconds.

  • result_reuse_enable – Enable Athena result reuse for this query.

  • result_reuse_minutes – Minutes to reuse cached results.

  • paramstyle – Parameter style (‘qmark’ or ‘pyformat’).

  • on_start_query_execution – Callback invoked with the query ID before execute() waits for the query: after the StartQueryExecution call, or after a reusable query ID is found through cache_size. Function signature: (query_id: str) -> None This allows early access to query_id for monitoring/cancellation.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index. For example: {"tags": "array(varchar)", "metadata": "map(varchar, integer)"}

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

Self reference for method chaining.

Example

>>> cursor.execute(
...     "SELECT * FROM table_with_complex_types",
...     result_set_type_hints={
...         "tags": "array(varchar)",
...         "metadata": "map(varchar, integer)",
...     }
... )
executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute a SQL query multiple times with different parameters.

On success, rowcount is the sum of the affected row counts, or -1 if any execution has an unknown count. An empty parameter list sets it to 0. On failure, it is -1; earlier executions are not rolled back. Result sets are discarded.

On failure, query_id retains the current query ID when available. If parameter iteration fails, this can identify the last successful execution.

Parameters:
  • operation – SQL query string to execute.

  • seq_of_parameters – Sequence of parameter sets, one per execution.

  • **kwargs – Additional keyword arguments passed to each execute().

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows from the result set.

Returns:

The remaining rows.

Raises:

ProgrammingError – If no result set is available.

fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch multiple rows from the result set.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The fetched rows.

Raises:

ProgrammingError – If no result set is available.

fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result set.

Returns:

The next row (a tuple, or a dict for dict cursors), or None if no more rows.

Raises:

ProgrammingError – If no result set is available.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

property has_result_set: bool

Whether the cursor has a result set.

property kms_key: str | None

The KMS key used to encrypt the query results.

list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

Asynchronous Cursors

class pyathena.async_cursor.AsyncCursor(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, max_workers: int = 20, arraysize: int = 1000, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs)[source]

Asynchronous cursor for non-blocking Athena query execution.

This cursor allows multiple queries to be executed concurrently without blocking the main thread. It’s useful for applications that need to execute multiple queries in parallel or perform other work while queries are running.

The cursor maintains a thread pool for executing queries asynchronously and provides methods to check query status and retrieve results when ready.

arraysize

Default number of rows that fetchmany() returns on the result sets this cursor creates.

Example

>>> cursor = connection.cursor(AsyncCursor)
>>>
>>> # Execute multiple queries concurrently
>>> query_id1, future1 = cursor.execute("SELECT COUNT(*) FROM table1")
>>> query_id2, future2 = cursor.execute("SELECT COUNT(*) FROM table2")
>>> query_id3, future3 = cursor.execute("SELECT COUNT(*) FROM table3")
>>>
>>> # Check if queries are done and get results
>>> if future1.done():
...     result1 = future1.result().fetchall()
>>>
>>> # Wait for all to complete
>>> results = [f.result().fetchall() for f in [future1, future2, future3]]

Note

Each execute() call returns a (query_id, future) tuple. The future resolves to the result set, which provides the column descriptions and the fetch methods. description(query_id) also returns the column descriptions as a Future.

__init__(s3_staging_dir: str | None = None, schema_name: str | None = None, catalog_name: str | None = None, work_group: str | None = None, poll_interval: float = 1, encryption_option: str | None = None, kms_key: str | None = None, kill_on_interrupt: bool = True, max_workers: int = 20, arraysize: int = 1000, result_reuse_enable: bool = False, result_reuse_minutes: int = 60, **kwargs) → None[source]

Initialize an AsyncCursor.

Parameters:
  • s3_staging_dir – S3 location for query results.

  • schema_name – Default schema name.

  • catalog_name – Default catalog name.

  • work_group – Athena workgroup name.

  • poll_interval – Query status polling interval in seconds.

  • encryption_option – S3 encryption option (SSE_S3, SSE_KMS, CSE_KMS).

  • kms_key – KMS key for encryption.

  • kill_on_interrupt – Cancel a query whose start in execute() is interrupted by KeyboardInterrupt. Waiting runs on worker threads, which do not receive the interrupt.

  • max_workers – Maximum number of threads in the cursor’s thread pool.

  • arraysize – Default number of rows per fetchmany() call of the result sets the cursor creates.

  • result_reuse_enable – Enable Athena query result reuse.

  • result_reuse_minutes – Maximum age in minutes of a reused result.

  • **kwargs – Arguments forwarded to BaseCursor.__init__, such as connection, converter, formatter, and retry_config.

Raises:

ProgrammingError – If arraysize is not between 1 and CursorIterator.DEFAULT_FETCH_SIZE.

property arraysize: int

The default number of rows per fetchmany() call of the result sets.

close(wait: bool = False) → None[source]

Close the cursor.

description(query_id: str) → Future[list[tuple[str, str, None, None, int, int, str]] | None][source]

Get the column descriptions of a query’s result set asynchronously.

The future waits for the query to finish before it reads the result set.

Parameters:

query_id – The Athena query execution ID.

Returns:

Future object containing the DB API 2.0 column descriptions, or None.

query_execution(query_id: str) → Future[AthenaQueryExecution][source]

Get query execution details asynchronously.

Retrieves the current execution status and metadata for a query. This is useful for monitoring query progress without blocking.

Parameters:

query_id – The Athena query execution ID.

Returns:

Future object containing AthenaQueryExecution with query details.

poll(query_id: str) → Future[AthenaQueryExecution][source]

Poll for query completion asynchronously.

Waits for the query to complete (succeed, fail, or be cancelled) and returns the final execution status. This method blocks until completion but runs the polling in a background thread.

Parameters:

query_id – The Athena query execution ID to poll.

Returns:

Future object containing the final AthenaQueryExecution status.

Note

This method performs polling internally, so it will take time proportional to your query execution duration.

execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → tuple[str, Future[AthenaResultSet | Any]][source]

Execute a SQL query asynchronously.

Starts query execution on Amazon Athena and returns immediately without waiting for completion. The query runs in the background while your application can continue with other work.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters (optional).

  • work_group – Athena workgroup to use (optional).

  • s3_staging_dir – S3 location for query results (optional).

  • cache_size – Number of queries to check for result caching (optional).

  • cache_expiration_time – Cache expiration time in seconds (optional).

  • result_reuse_enable – Enable result reuse for identical queries (optional).

  • result_reuse_minutes – Result reuse duration in minutes (optional).

  • paramstyle – Parameter style to use (optional).

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

  • query_id: Athena query execution ID for tracking

  • future: Future object for result retrieval

Return type:

Tuple of (query_id, future) where

Example

>>> query_id, future = cursor.execute("SELECT * FROM large_table")
>>> print(f"Query started: {query_id}")
>>> # Do other work while query runs...
>>> result_set = future.result()  # Wait for completion
executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None[source]

Execute multiple queries asynchronously (not supported).

This method is not supported for asynchronous cursors because managing multiple concurrent queries would be complex and resource-intensive.

Parameters:
  • operation – SQL query string.

  • seq_of_parameters – Sequence of parameter sets.

  • **kwargs – Additional arguments.

Raises:

NotSupportedError – Always raised as this operation is not supported.

Note

For bulk operations, consider using execute() with parameterized queries or batch processing patterns instead.

cancel(query_id: str) → Future[None][source]

Cancel a running query asynchronously.

Submits a cancellation request for the specified query. The cancellation itself runs asynchronously in the background.

Parameters:

query_id – The Athena query execution ID to cancel.

Returns:

Future object that completes when the cancellation request finishes.

Example

>>> query_id, future = cursor.execute("SELECT * FROM huge_table")
>>> # Later, cancel the query
>>> cancel_future = cursor.cancel(query_id)
>>> cancel_future.result()  # Wait for cancellation to complete
LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
property connection: Connection[Any]

The connection that created this cursor.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

class pyathena.async_cursor.AsyncDictCursor(dict_type: type[Any] | None = None, **kwargs)[source]

Asynchronous cursor that returns query results as dictionaries.

Combines the asynchronous execution capabilities of AsyncCursor with the dictionary-based result format of DictCursor. Results are returned as dictionaries where column names are keys, making it easier to access column values by name rather than position.

Example

>>> cursor = connection.cursor(AsyncDictCursor)
>>> query_id, future = cursor.execute("SELECT id, name, email FROM users")
>>> result_set = future.result()
>>> row = result_set.fetchone()
>>> print(f"User: {row['name']} ({row['email']})")
__init__(dict_type: type[Any] | None = None, **kwargs) → None[source]

Initialize an AsyncDictCursor.

Parameters:
  • dict_type – The type used to build each row of this cursor’s result sets. If None, the result set class’s dict_type is used.

  • **kwargs – Arguments forwarded to AsyncCursor.__init__.

LIST_DATABASES_MAX_RESULTS = 50
LIST_QUERY_EXECUTIONS_MAX_RESULTS = 50
LIST_TABLE_METADATA_MAX_RESULTS = 50
property arraysize: int

The default number of rows per fetchmany() call of the result sets.

cancel(query_id: str) → Future[None]

Cancel a running query asynchronously.

Submits a cancellation request for the specified query. The cancellation itself runs asynchronously in the background.

Parameters:

query_id – The Athena query execution ID to cancel.

Returns:

Future object that completes when the cancellation request finishes.

Example

>>> query_id, future = cursor.execute("SELECT * FROM huge_table")
>>> # Later, cancel the query
>>> cancel_future = cursor.cancel(query_id)
>>> cancel_future.result()  # Wait for cancellation to complete
close(wait: bool = False) → None

Close the cursor.

property connection: Connection[Any]

The connection that created this cursor.

description(query_id: str) → Future[list[tuple[str, str, None, None, int, int, str]] | None]

Get the column descriptions of a query’s result set asynchronously.

The future waits for the query to finish before it reads the result set.

Parameters:

query_id – The Athena query execution ID.

Returns:

Future object containing the DB API 2.0 column descriptions, or None.

execute(operation: str, parameters: dict[str, Any] | list[str] | None = None, work_group: str | None = None, s3_staging_dir: str | None = None, cache_size: int | None = None, cache_expiration_time: int | None = None, result_reuse_enable: bool | None = None, result_reuse_minutes: int | None = None, paramstyle: str | None = None, result_set_type_hints: dict[str | int, str] | None = None, *, options: ExecuteOptions | None = None, **kwargs) → tuple[str, Future[AthenaResultSet | Any]]

Execute a SQL query asynchronously.

Starts query execution on Amazon Athena and returns immediately without waiting for completion. The query runs in the background while your application can continue with other work.

Parameters:
  • operation – SQL query string to execute.

  • parameters – Query parameters (optional).

  • work_group – Athena workgroup to use (optional).

  • s3_staging_dir – S3 location for query results (optional).

  • cache_size – Number of queries to check for result caching (optional).

  • cache_expiration_time – Cache expiration time in seconds (optional).

  • result_reuse_enable – Enable result reuse for identical queries (optional).

  • result_reuse_minutes – Result reuse duration in minutes (optional).

  • paramstyle – Parameter style to use (optional).

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

  • options – Shared execution options as an ExecuteOptions instance. Individual keyword arguments take precedence over options fields.

  • **kwargs – Additional execution parameters.

Returns:

  • query_id: Athena query execution ID for tracking

  • future: Future object for result retrieval

Return type:

Tuple of (query_id, future) where

Example

>>> query_id, future = cursor.execute("SELECT * FROM large_table")
>>> print(f"Query started: {query_id}")
>>> # Do other work while query runs...
>>> result_set = future.result()  # Wait for completion
executemany(operation: str, seq_of_parameters: list[dict[str, Any] | list[str] | None], **kwargs) → None

Execute multiple queries asynchronously (not supported).

This method is not supported for asynchronous cursors because managing multiple concurrent queries would be complex and resource-intensive.

Parameters:
  • operation – SQL query string.

  • seq_of_parameters – Sequence of parameter sets.

  • **kwargs – Additional arguments.

Raises:

NotSupportedError – Always raised as this operation is not supported.

Note

For bulk operations, consider using execute() with parameterized queries or batch processing patterns instead.

static get_default_converter(unload: bool = False) → DefaultTypeConverter | Any

Get the default type converter for this cursor class.

Parameters:

unload – Whether the converter is for UNLOAD operations. Some cursor types may return different converters for UNLOAD operations.

Returns:

The default type converter instance for this cursor type.

get_table_metadata(table_name: str, catalog_name: str | None = None, schema_name: str | None = None, logging_: bool = True) → AthenaTableMetadata

Get one table’s metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • table_name – The table name.

  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • logging – Whether to log a failed request.

Returns:

The table’s metadata.

Raises:

OperationalError – If the request fails, including when the table does not exist.

list_databases(catalog_name: str | None, max_results: int | None = None) → list[AthenaDatabase]

List the catalog’s databases.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • max_results – The page size of each request.

Returns:

The catalog’s databases.

Raises:

OperationalError – If the request fails.

list_table_metadata(catalog_name: str | None = None, schema_name: str | None = None, expression: str | None = None, max_results: int | None = None, logging_: bool = True) → list[AthenaTableMetadata]

List a database’s table metadata.

In AwsDataCatalog and S3 Tables catalogs, a throttled request is answered from the AWS Glue Data Catalog; see glue_metadata_fallback.

Parameters:
  • catalog_name – The catalog, or None for the cursor’s catalog.

  • schema_name – The database, or None for the cursor’s schema.

  • expression – A table name pattern.

  • max_results – The page size of each request.

  • logging – Whether to log a failed request.

Returns:

The metadata of the database’s tables.

Raises:

OperationalError – If the request fails.

poll(query_id: str) → Future[AthenaQueryExecution]

Poll for query completion asynchronously.

Waits for the query to complete (succeed, fail, or be cancelled) and returns the final execution status. This method blocks until completion but runs the polling in a background thread.

Parameters:

query_id – The Athena query execution ID to poll.

Returns:

Future object containing the final AthenaQueryExecution status.

Note

This method performs polling internally, so it will take time proportional to your query execution duration.

query_execution(query_id: str) → Future[AthenaQueryExecution]

Get query execution details asynchronously.

Retrieves the current execution status and metadata for a query. This is useful for monitoring query progress without blocking.

Parameters:

query_id – The Athena query execution ID.

Returns:

Future object containing AthenaQueryExecution with query details.

setinputsizes(sizes)

Accept input sizes as DB API 2.0 requires, and ignore them.

Parameters:

sizes – Sequence of parameter types or sizes.

setoutputsize(size, column=None)

Accept a column buffer size as DB API 2.0 requires, and ignore it.

Parameters:
  • size – Buffer size for large columns.

  • column – Index of the column the size applies to, or None for all large columns.

Result Sets

class pyathena.result_set.AthenaResultSet(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, _pre_fetch: bool = True, result_set_type_hints: dict[str | int, str] | None = None)[source]

Result set for Athena query execution using the GetQueryResults API.

This class provides a DB API 2.0 compliant result set implementation that fetches query results from Amazon Athena. It uses the GetQueryResults API to retrieve data in paginated chunks, converting each value according to its Athena data type.

The result set exposes query execution metadata (timing, data scanned, state, etc.) through read-only properties, allowing inspection of query performance and status.

This is the base result set implementation used by the standard Cursor. Specialized implementations exist for different output formats:

Example

>>> cursor.execute("SELECT * FROM my_table")
>>> result_set = cursor.result_set
>>> print(f"Query ID: {result_set.query_id}")
>>> print(f"Data scanned: {result_set.data_scanned_in_bytes} bytes")
>>> for row in result_set:
...     print(row)
__init__(connection: Connection[Any], converter: Converter, query_execution: AthenaQueryExecution, arraysize: int, retry_config: RetryConfig, _pre_fetch: bool = True, result_set_type_hints: dict[str | int, str] | None = None) → None[source]

Initialize the result set and fetch the first page if the query succeeded.

Parameters:
  • connection – The connection that ran the query.

  • converter – The converter for result values.

  • query_execution – The query execution whose results to read.

  • arraysize – The number of rows per GetQueryResults page and the default fetchmany() size.

  • retry_config – The retry configuration for API calls.

  • _pre_fetch – Whether to fetch the first page here when the query succeeded. The async result set passes False and fetches it itself.

  • result_set_type_hints – Athena type signatures for complex-type columns, keyed by column name (case-insensitive) or zero-based column index.

Raises:
property database: str | None

The database in the QueryExecutionContext of the query.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

property query_id: str | None

The ID of the query execution.

property query: str | None

The SQL statement that the query execution ran.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property work_group: str | None

The work group in which the query ran.

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property completion_date_time: datetime | None

The date and time when the query completed.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_type: int | None

The ErrorType code of the query failure.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property error_message: str | None

The ErrorMessage that describes the query failure.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

property output_location: str | None

The S3 location of the query results.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property is_unload: bool

Check if the query is an UNLOAD statement.

Returns:

True if the query is an UNLOAD statement, False otherwise.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property kms_key: str | None

The KMS key used to encrypt the query results.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions.

None without result metadata, or for INSERT, UPDATE, DELETE, and MERGE.

property connection: Connection[Any]

The connection of the result set; raises ProgrammingError if closed.

fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None[source]

Fetch the next row of the result.

fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]][source]

Fetch the next set of rows of the query result.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The rows, fewer than size when the result is exhausted.

fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]][source]

Fetch all remaining rows of the query result.

Returns:

The remaining rows.

property is_closed: bool

Whether the result set is closed.

close() → None[source]

Close the result set and discard its query execution, metadata, and rows.

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
property arraysize: int

The default number of rows per fetchmany() call.

property rowcount: int

The number of rows affected by the last operation, or -1 if it is unknown.

property rownumber: int | None

The zero-based index of the next row, or None if it is unknown.

class pyathena.result_set.AthenaDictResultSet(*args: Any, dict_type: type[Any] | None = None, **kwargs: Any)[source]

A result set that returns each row as a dictionary keyed by column name.

__init__(*args: Any, dict_type: type[Any] | None = None, **kwargs: Any) → None[source]

Initialize the result set with an optional row type for this instance.

Parameters:
  • *args – Positional arguments passed to the next __init__ in the MRO.

  • dict_type – The type used to build each row of this result set. If None, the class attribute dict_type is used.

  • **kwargs – Keyword arguments passed to the next __init__ in the MRO.

dict_type

alias of dict

DEFAULT_FETCH_SIZE: int = 1000
DEFAULT_RESULT_REUSE_MINUTES = 60
property arraysize: int

The default number of rows per fetchmany() call.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

close() → None

Close the result set and discard its query execution, metadata, and rows.

property completion_date_time: datetime | None

The date and time when the query completed.

property connection: Connection[Any]

The connection of the result set; raises ProgrammingError if closed.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property database: str | None

The database in the QueryExecutionContext of the query.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions.

None without result metadata, or for INSERT, UPDATE, DELETE, and MERGE.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_message: str | None

The ErrorMessage that describes the query failure.

property error_type: int | None

The ErrorType code of the query failure.

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

fetchall() → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch all remaining rows of the query result.

Returns:

The remaining rows.

fetchmany(size: int | None = None) → list[tuple[Any | None, ...] | dict[Any, Any | None]]

Fetch the next set of rows of the query result.

Parameters:

size – Maximum number of rows to fetch. If None or not positive, arraysize is used.

Returns:

The rows, fewer than size when the result is exhausted.

fetchone() → tuple[Any | None, ...] | dict[Any, Any | None] | None

Fetch the next row of the result.

property is_closed: bool

Whether the result set is closed.

property is_unload: bool

Check if the query is an UNLOAD statement.

Returns:

True if the query is an UNLOAD statement, False otherwise.

property kms_key: str | None

The KMS key used to encrypt the query results.

property output_location: str | None

The S3 location of the query results.

property query: str | None

The SQL statement that the query execution ran.

property query_id: str | None

The ID of the query execution.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property rowcount: int

The number of rows affected by the last operation, or -1 if it is unknown.

property rownumber: int | None

The zero-based index of the next row, or None if it is unknown.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property work_group: str | None

The work group in which the query ran.

class pyathena.result_set.WithResultSet(arraysize: int | None = None, **kwargs)[source]

Mixin that keeps a cursor’s query ID and result set.

Provides the query ID, the result set and its properties, arraysize, rownumber, rowcount, and close. WithFetch and WithAsyncFetch list it before BaseCursor / AioBaseCursor and CursorIterator, so that these members take precedence over theirs.

__init__(arraysize: int | None = None, **kwargs) → None[source]

Initialize the cursor with no query ID and no result set.

Parameters:
  • arraysize – Default number of rows per fetchmany() call, validated by the arraysize setter. If None, DEFAULT_FETCH_SIZE is used.

  • **kwargs – Arguments passed to the next __init__ in the MRO.

Raises:

ProgrammingError – If arraysize is outside the range the cursor’s arraysize setter accepts.

property result_set: AthenaResultSet | None

The result set of the last executed query.

Returns:

The result set, or None before a query succeeds or after a reset.

property has_result_set: bool

Whether the cursor has a result set.

property description: list[tuple[str, str, None, None, int, int, str]] | None

The DB API 2.0 column descriptions of the result set, or None without one.

property database: str | None

The database in the QueryExecutionContext of the query.

property catalog: str | None

The data catalog in the QueryExecutionContext of the query.

property query_id: str | None

The query execution ID of the last execution.

With cache_size or cache_expiration_time, this can be the ID of a previous execution whose result is reused.

Returns:

The query execution ID, or None if there is none since the last reset.

property query: str | None

The SQL statement that the query execution ran.

property statement_type: str | None

The StatementType of the query, such as DDL, DML, or UTILITY.

property substatement_type: str | None

The SubstatementType of the query, such as INSERT or MERGE.

property work_group: str | None

The work group in which the query ran.

property execution_parameters: list[str]

The ExecutionParameters values of the query.

property state: str | None

The state of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason that gives further detail about the state.

property submission_date_time: datetime | None

The date and time when the query was submitted.

property completion_date_time: datetime | None

The date and time when the query completed.

property error_category: int | None

1 for system, 2 for user, 3 for other.

Type:

The ErrorCategory of the failure

property error_type: int | None

The ErrorType code of the query failure.

property retryable: bool | None

Whether Athena reports the query failure as retryable.

property error_message: str | None

The ErrorMessage that describes the query failure.

property data_scanned_in_bytes: int | None

The number of bytes that the query scanned.

property engine_execution_time_in_millis: int | None

The time in milliseconds that the query engine took to run the query.

property query_queue_time_in_millis: int | None

The time in milliseconds that the query waited in the queue.

property total_execution_time_in_millis: int | None

The total time in milliseconds that Athena took to run the query.

property query_planning_time_in_millis: int | None

The time in milliseconds that Athena took to plan the query.

property service_processing_time_in_millis: int | None

The time in milliseconds that Athena took to publish the query results.

property output_location: str | None

The S3 location of the query results.

property data_manifest_location: str | None

The S3 location of the data manifest that lists the files the query wrote.

property reused_previous_result: bool | None

Whether Athena reused a previous query result instead of running the query.

property encryption_option: str | None

The EncryptionOption of the query results, such as SSE_S3 or SSE_KMS.

property kms_key: str | None

The KMS key used to encrypt the query results.

property expected_bucket_owner: str | None

The AWS account ID expected to own the S3 bucket of the query results.

property s3_acl_option: str | None

The S3AclOption of the query results, such as BUCKET_OWNER_FULL_CONTROL.

property selected_engine_version: str | None

The Athena engine version selected to run the query.

property effective_engine_version: str | None

The Athena engine version that ran the query.

property result_reuse_enabled: bool | None

Whether reuse of previous query results by age is enabled for the query.

property result_reuse_minutes: int | None

The maximum age in minutes of a previous query result that Athena can reuse.

property rowcount: int

Get the number of rows affected by the last operation.

For SELECT statements, this returns -1 as per DB API 2.0 specification. For DML operations (INSERT, UPDATE, DELETE) and CTAS, this returns the number of affected rows. After a successful executemany(), this is the sum across executions, or -1 if any count is unknown.

Returns:

The number of rows, or -1 if not applicable or unknown.

property arraysize: int

The default number of rows per fetchmany() call.

execute() passes it to the new result set, so a change applies to the result sets of later executions. Setting it to zero or a negative value raises ProgrammingError.

Returns:

The default number of rows per fetchmany() call.

property rownumber: int | None

The zero-based index of the next row in the result set.

Returns:

The row index, or None if there is no result set or the index is unknown.

close() → None[source]

Close the cursor and release associated resources.