READ_PARQUET
READ_PARQUET is a table function: it reads a Parquet file (or a set of files matched
by a glob) directly by path, without registering it as a table in a catalog first. Use
it in a FROM clause wherever a table name is expected.
Syntax
FROM READ_PARQUET(<path>)Parameters
<path>— single string literal giving the file path (or glob pattern matching multiple files) to read.
READ_PARQUET takes no other arguments — Parquet's schema is read straight from the
file's own footer, so there is nothing to configure the way there is for
READ_CSV/READ_JSONL.
Examples
Query a Single File
SELECT *
FROM READ_PARQUET('data/packages.parquet');Query a Remote File
SELECT *
FROM READ_PARQUET('https://example.com/data/packages.parquet');Query a Set of Files with a Glob
SELECT *
FROM READ_PARQUET('data/packages-*.parquet');Use Inside CREATE TABLE AS
CREATE TABLE my_workspace.my_collection.packages AS
SELECT *
FROM READ_PARQUET('https://example.com/data/packages-*.parquet');Alias the Relation
SELECT p.name
FROM READ_PARQUET('data/packages.parquet') AS p
WHERE p.active = TRUE;Notes
- Column names and types come directly from the schema embedded in the Parquet file(s);
there is no
AS alias(col1, col2, ...)form to rename columns — useSELECT ... AS new_nameinstead. A plain relation alias (AS alias, no column list) is supported. - Standard filter and column pushdown apply:
WHEREpredicates and the columns your query actually references are pushed into the scan. - A glob path (containing
*,?, or[) matches multiple files; their combined content is read as one relation. Non-.parquetfiles matched by a glob are silently excluded. gs://bucket/objectpaths are supported and are always fetched anonymously (a public GCS object is read; a private one fails with an error) — Opteryx never signs a request or uses platform credentials on your behalf for a path given toREAD_PARQUET. Glob patterns are not supported forgs://paths, because listing a bucket's contents needs a permission a public, unauthenticated read does not have. Usegs://, notgcs://.