Data Models

This section covers Athena query execution models and configuration classes.

Query Execution

class pyathena.model.AthenaQueryExecution(response: dict[str, Any])[source]

Represents an Athena query execution with status and metadata.

This class encapsulates information about a query execution in Amazon Athena, including its current state, statistics, error information, and result metadata. It’s primarily used internally by PyAthena cursors but can be useful for monitoring and debugging query execution.

Query States:
  • QUEUED: Query is waiting to be executed

  • RUNNING: Query is currently executing

  • SUCCEEDED: Query completed successfully

  • FAILED: Query execution failed

  • CANCELLED: Query was cancelled

Statement Types:
  • DDL: Data Definition Language (CREATE, DROP, ALTER)

  • DML: Data Manipulation Language (SELECT, INSERT, UPDATE, DELETE)

  • UTILITY: Utility statements (SHOW, DESCRIBE, EXPLAIN)

Example

>>> # AsyncCursor returns the query execution through a Future
>>> query_id, future = cursor.execute("SELECT COUNT(*) FROM my_table")
>>> query_execution = cursor.query_execution(query_id).result()
>>> print(f"Query ID: {query_execution.query_id}")
>>> print(f"State: {query_execution.state}")
>>> print(f"Data scanned: {query_execution.data_scanned_in_bytes} bytes")

See also

AWS Athena QueryExecution API reference: https://docs.aws.amazon.com/athena/latest/APIReference/API_QueryExecution.html

STATE_QUEUED: str = 'QUEUED'
STATE_RUNNING: str = 'RUNNING'
STATE_SUCCEEDED: str = 'SUCCEEDED'
STATE_FAILED: str = 'FAILED'
STATE_CANCELLED: str = 'CANCELLED'
TERMINAL_STATES: tuple[str, ...] = ('SUCCEEDED', 'FAILED', 'CANCELLED')
STATEMENT_TYPE_DDL: str = 'DDL'
STATEMENT_TYPE_DML: str = 'DML'
STATEMENT_TYPE_UTILITY: str = 'UTILITY'
ENCRYPTION_OPTION_SSE_S3: str = 'SSE_S3'
ENCRYPTION_OPTION_SSE_KMS: str = 'SSE_KMS'
ENCRYPTION_OPTION_CSE_KMS: str = 'CSE_KMS'
ERROR_CATEGORY_SYSTEM: int = 1
ERROR_CATEGORY_USER: int = 2
ERROR_CATEGORY_OTHER: int = 3
S3_ACL_OPTION_BUCKET_OWNER_FULL_CONTROL = 'BUCKET_OWNER_FULL_CONTROL'
__init__(response: dict[str, Any]) → None[source]

Initialize the query execution from a GetQueryExecution response.

Parameters:

response – The API response containing a QueryExecution object.

Raises:

DataError – If QueryExecution, QueryExecutionId, Query, or Status is missing from the response.

property database: str | None

The Database of the query execution context.

property catalog: str | None

The Catalog of the query execution context.

property query_id: str | None

The QueryExecutionId of the query.

property query: str | None

The Query string that was executed.

property statement_type: str | None

The StatementType of the query, such as DDL or DML.

property substatement_type: str | None

The SubstatementType of the query.

property work_group: str | None

The WorkGroup in which the query ran.

property execution_parameters: list[str]

The ExecutionParameters of the query, or an empty list.

property state: str | None

The State of the query execution, such as RUNNING or SUCCEEDED.

property state_change_reason: str | None

The StateChangeReason of the query execution status.

property submission_date_time: datetime | None

The SubmissionDateTime of the query.

property completion_date_time: datetime | None

The CompletionDateTime of the query.

property error_category: int | None

The ErrorCategory of the AthenaError.

property error_type: int | None

The ErrorType of the AthenaError.

property retryable: bool | None

The Retryable flag of the AthenaError.

property error_message: str | None

The ErrorMessage of the AthenaError.

property data_scanned_in_bytes: int | None

The DataScannedInBytes statistic of the query.

property engine_execution_time_in_millis: int | None

The EngineExecutionTimeInMillis statistic of the query.

property query_queue_time_in_millis: int | None

The QueryQueueTimeInMillis statistic of the query.

property total_execution_time_in_millis: int | None

The TotalExecutionTimeInMillis statistic of the query.

property query_planning_time_in_millis: int | None

The QueryPlanningTimeInMillis statistic of the query.

property service_pre_processing_time_in_millis: int | None

The ServicePreProcessingTimeInMillis statistic of the query.

property service_processing_time_in_millis: int | None

The ServiceProcessingTimeInMillis statistic of the query.

property dpu_count: float | None

The DpuCount statistic of the query.

property output_location: str | None

The OutputLocation of the result configuration.

property data_manifest_location: str | None

The DataManifestLocation statistic of the query.

property reused_previous_result: bool | None

The ReusedPreviousResult flag of the result reuse information.

property encryption_option: str | None

The EncryptionOption of the result encryption configuration.

property kms_key: str | None

The KmsKey of the result encryption configuration.

property expected_bucket_owner: str | None

The ExpectedBucketOwner of the result configuration.

property s3_acl_option: str | None

The S3AclOption of the result ACL configuration.

property selected_engine_version: str | None

The SelectedEngineVersion of the query.

property effective_engine_version: str | None

The EffectiveEngineVersion of the query.

property result_reuse_enabled: bool | None

The Enabled flag of the result reuse by age configuration.

property result_reuse_minutes: int | None

The MaxAgeInMinutes of the result reuse by age configuration.

property managed_query_results_enabled: bool | None

The Enabled flag of the managed query results configuration.

property managed_query_results_kms_key: str | None

The KmsKey of the managed query results encryption configuration.

property enable_s3_access_grants: bool | None

The EnableS3AccessGrants flag of the S3 Access Grants configuration.

property create_user_level_prefix: bool | None

The CreateUserLevelPrefix flag of the S3 Access Grants configuration.

property s3_access_grants_authentication_type: str | None

The AuthenticationType of the S3 Access Grants configuration.

class pyathena.model.AthenaCalculationExecution(response: dict[str, Any])[source]

Represents a complete Athena calculation execution with status and results.

This class extends AthenaCalculationExecutionStatus to include additional information about the calculation execution, including session details, working directory, and result locations in S3.

Attributes are inherited from AthenaCalculationExecutionStatus for state and timing information.

See also

AWS Athena GetCalculationExecution API reference: https://docs.aws.amazon.com/athena/latest/APIReference/API_GetCalculationExecution.html

__init__(response: dict[str, Any]) → None[source]

Initialize the calculation execution from a GetCalculationExecution response.

Parameters:

response – The API response containing the calculation fields, Status, Statistics, and an optional Result object.

Raises:

DataError – If Status, Statistics, CalculationExecutionId, or SessionId is missing from the response.

property calculation_id: str | None

The CalculationExecutionId of the calculation.

property session_id: str | None

The SessionId of the session that ran the calculation.

property description: str | None

The Description of the calculation.

property working_directory: str | None

The WorkingDirectory of the calculation.

property std_out_s3_uri: str | None

The StdOutS3Uri of the calculation result.

property std_error_s3_uri: str | None

The StdErrorS3Uri of the calculation result.

property result_s3_uri: str | None

The ResultS3Uri of the calculation result.

property result_type: str | None

The ResultType of the calculation result.

class pyathena.model.AthenaCalculationExecutionStatus(response: dict[str, Any])[source]

Status information for an Athena calculation execution.

This class represents the current state and statistics of a calculation execution in Amazon Athena’s notebook or interactive session environment. It tracks the calculation’s lifecycle from creation through completion.

Calculation States:
  • CREATING: Calculation is being created

  • CREATED: Calculation has been created

  • QUEUED: Calculation is waiting to execute

  • RUNNING: Calculation is currently executing

  • CANCELING: Calculation is being cancelled

  • CANCELED: Calculation was cancelled

  • COMPLETED: Calculation completed successfully

  • FAILED: Calculation execution failed

See also

AWS Athena CalculationExecutionStatus API reference: https://docs.aws.amazon.com/athena/latest/APIReference/API_CalculationStatus.html

STATE_CREATING: str = 'CREATING'
STATE_CREATED: str = 'CREATED'
STATE_QUEUED: str = 'QUEUED'
STATE_RUNNING: str = 'RUNNING'
STATE_CANCELING: str = 'CANCELING'
STATE_CANCELED: str = 'CANCELED'
STATE_COMPLETED: str = 'COMPLETED'
STATE_FAILED: str = 'FAILED'
TERMINAL_STATES: tuple[str, ...] = ('COMPLETED', 'FAILED', 'CANCELED')
__init__(response: dict[str, Any]) → None[source]

Initialize the calculation status from an Athena API response.

Parameters:

response – The API response containing Status and Statistics objects.

Raises:

DataError – If Status or Statistics is missing from the response.

property state: str | None

The State of the calculation, such as RUNNING or COMPLETED.

property state_change_reason: str | None

The StateChangeReason of the calculation status.

property submission_date_time: datetime | None

The SubmissionDateTime of the calculation.

property completion_date_time: datetime | None

The CompletionDateTime of the calculation.

property dpu_execution_in_millis: int | None

The DpuExecutionInMillis statistic of the calculation.

property progress: str | None

The Progress statistic of the calculation.

Session Management

class pyathena.model.AthenaSessionStatus(response: dict[str, Any])[source]

Status information for an Athena interactive session.

This class represents the current state of an interactive session in Amazon Athena, used for notebook and Spark workloads. Sessions provide a persistent environment for running multiple calculations.

Session States:
  • CREATING: Session is being created

  • CREATED: Session has been created

  • IDLE: Session is idle and ready for calculations

  • BUSY: Session is executing a calculation

  • TERMINATING: Session is being terminated

  • TERMINATED: Session has been terminated

  • DEGRADED: Session is in a degraded state

  • FAILED: Session creation or execution failed

STATE_CREATING: str = 'CREATING'
STATE_CREATED: str = 'CREATED'
STATE_IDLE: str = 'IDLE'
STATE_BUSY: str = 'BUSY'
STATE_TERMINATING: str = 'TERMINATING'
STATE_TERMINATED: str = 'TERMINATED'
STATE_DEGRADED: str = 'DEGRADED'
STATE_FAILED: str = 'FAILED'
__init__(response: dict[str, Any]) → None[source]

Initialize the session status from an Athena API response.

Parameters:

response – The API response containing SessionId and a Status object.

Raises:

DataError – If Status is missing from the response.

property session_id: str | None

The SessionId of the session.

property state: str | None

The State of the session, such as IDLE or BUSY.

property state_change_reason: str | None

The StateChangeReason of the session status.

property start_date_time: datetime | None

The StartDateTime of the session.

property last_modified_date_time: datetime | None

The LastModifiedDateTime of the session.

property end_date_time: datetime | None

The EndDateTime of the session.

property idle_since_date_time: datetime | None

The IdleSinceDateTime of the session.

Database and Table Metadata

class pyathena.model.AthenaDatabase(response)[source]

Represents an Athena database (schema) and its metadata.

This class encapsulates information about a database in the AWS Glue Data Catalog that is accessible through Amazon Athena. Databases serve as containers for tables and views.

__init__(response)[source]

Initialize the database from an Athena API response.

Parameters:

response – A dictionary containing a Database object.

Raises:

DataError – If Database is missing from the response.

property name: str | None

The Name of the database.

property description: str | None

The Description of the database.

property parameters: dict[str, str]

The Parameters of the database, or an empty dictionary.

class pyathena.model.AthenaTableMetadata(response)[source]

Represents comprehensive metadata for an Athena table.

This class contains detailed information about a table in the AWS Glue Data Catalog, including columns, partition keys, storage format, serialization library, and various table properties.

The class provides convenient properties for accessing common table attributes like location, file format, compression, and SerDe configuration.

__init__(response)[source]

Initialize the table metadata from an Athena API response.

Parameters:

response – A dictionary containing a TableMetadata object.

Raises:

DataError – If TableMetadata is missing from the response.

property name: str | None

The Name of the table.

property create_time: datetime | None

The CreateTime of the table.

property last_access_time: datetime | None

The LastAccessTime of the table.

property table_type: str | None

The TableType of the table.

property columns: list[AthenaTableMetadataColumn]

The Columns of the table.

property partition_keys: list[AthenaTableMetadataPartitionKey]

The PartitionKeys of the table.

property parameters: dict[str, str]

The Parameters of the table, or an empty dictionary.

property comment: str | None

The comment table parameter.

property location: str | None

The location table parameter.

property input_format: str | None

The inputformat table parameter.

property output_format: str | None

The outputformat table parameter.

property row_format: str | None

The SERDE '<lib>' clause built from serde_serialization_lib, or None.

property file_format: str | None

The INPUTFORMAT '...' OUTPUTFORMAT '...' clause, or None unless both are set.

property serde_serialization_lib: str | None

The serde.serialization.lib table parameter.

property compression: str | None

The compression codec from the table parameters, or None.

The first parameter present is used, in the order write.compression, serde.param.write.compression, parquet.compress, and orc.compress.

property serde_properties: dict[str, str]

The serde.param.-prefixed table parameters with the prefix removed.

property table_properties: dict[str, str]

The table parameters that do not start with serde.param..

File Formats and Compression

class pyathena.model.AthenaFileFormat[source]

Constants and utilities for Athena supported file formats.

This class provides constants for file formats supported by Amazon Athena and utility methods to check format types. These are commonly used when creating tables or configuring UNLOAD operations.

Supported formats:
  • SEQUENCEFILE: Hadoop SequenceFile format

  • TEXTFILE: Plain text files (default)

  • RCFILE: Record Columnar File format

  • ORC: Optimized Row Columnar format

  • PARQUET: Apache Parquet columnar format

  • AVRO: Apache Avro format

  • ION: Amazon Ion format

Example

>>> from pyathena.model import AthenaFileFormat
>>>
>>> # Check if format is Parquet
>>> if AthenaFileFormat.is_parquet("PARQUET"):
...     print("Using columnar format")
>>>
>>> # Use in UNLOAD operations
>>> format_type = AthenaFileFormat.FILE_FORMAT_PARQUET
>>> sql = f"UNLOAD (...) TO 's3://bucket/path/' WITH (format = '{format_type}')"
>>> cursor.execute(sql)

See also

AWS Documentation on supported file formats: https://docs.aws.amazon.com/athena/latest/ug/supported-serdes.html

FILE_FORMAT_SEQUENCEFILE: str = 'SEQUENCEFILE'
FILE_FORMAT_TEXTFILE: str = 'TEXTFILE'
FILE_FORMAT_RCFILE: str = 'RCFILE'
FILE_FORMAT_ORC: str = 'ORC'
FILE_FORMAT_PARQUET: str = 'PARQUET'
FILE_FORMAT_AVRO: str = 'AVRO'
FILE_FORMAT_ION: str = 'ION'
static is_parquet(value: str) → bool[source]

Check whether a file format name is PARQUET, ignoring case.

Parameters:

value – The file format name.

Returns:

True if the value is PARQUET, False otherwise.

static is_orc(value: str) → bool[source]

Check whether a file format name is ORC, ignoring case.

Parameters:

value – The file format name.

Returns:

True if the value is ORC, False otherwise.

class pyathena.model.AthenaCompression[source]

Constants and utilities for Athena supported compression formats.

This class provides constants for compression formats supported by Amazon Athena and utility methods to validate compression types. These are commonly used when creating tables, configuring UNLOAD operations, or optimizing data storage.

Supported compression formats:
  • BZIP2: BZIP2 compression

  • DEFLATE: DEFLATE compression

  • GZIP: GZIP compression (most common)

  • LZ4: LZ4 fast compression

  • LZO: LZO compression

  • SNAPPY: Snappy compression (good for Parquet)

  • ZLIB: ZLIB compression

  • ZSTD: Zstandard compression

Example

>>> from pyathena.model import AthenaCompression
>>>
>>> # Validate compression format
>>> if AthenaCompression.is_valid("GZIP"):
...     print("Valid compression format")
>>>
>>> # Use in UNLOAD operations
>>> compression = AthenaCompression.COMPRESSION_GZIP
>>> sql = f"UNLOAD (...) TO 's3://bucket/path/' WITH (compression = '{compression}')"
>>> cursor.execute(sql)

See also

AWS Documentation on compression formats: https://docs.aws.amazon.com/athena/latest/ug/compression-formats.html

Best practices for data compression in Athena: https://docs.aws.amazon.com/athena/latest/ug/compression-support.html

COMPRESSION_BZIP2: str = 'BZIP2'
COMPRESSION_DEFLATE: str = 'DEFLATE'
COMPRESSION_GZIP: str = 'GZIP'
COMPRESSION_LZ4: str = 'LZ4'
COMPRESSION_LZO: str = 'LZO'
COMPRESSION_SNAPPY: str = 'SNAPPY'
COMPRESSION_ZLIB: str = 'ZLIB'
COMPRESSION_ZSTD: str = 'ZSTD'
static is_valid(value: str) → bool[source]

Check whether a value is a supported compression format, ignoring case.

Parameters:

value – The compression format name.

Returns:

True if the value matches one of the COMPRESSION_* constants, False otherwise.