Wapgee Logowapgee

Parquet to CSV

API

Drop a Parquet file to see its schema, row count, and row groups, then preview the first rows and export CSV or JSON. Nothing is uploaded.

Drop a .parquet file, or choose one. Nothing is uploaded.

API

Call this tool over HTTP with a personal API token. The tool above keeps running in your browser. The API is a separate, server-side path. See API documentation for tokens, authentication, limits, and errors shared by every endpoint.

Read a Parquet file from a public link or a public S3 object and get CSV or JSON back. Formatting matches the tool page.

POST /api/v1/tools/parquet-to-csv

Daily allowance: 10 successful requests, shared by every token on your account. Resets at midnight UTC.

The file must be readable without credentials: a public https URL, or s3://bucket/key for a public object. The file is limited to 25 MB. Uncompressed and Snappy files are supported. Only the first limit rows are returned (default 10,000, maximum 100,000), and truncated says whether rows were left out. Private, loopback, and link-local addresses are blocked, including after redirects.

Request body

FieldDescription
urlRequired. A public http(s) URL or an s3://bucket/key address.
regionAWS region for an s3:// address, such as eu-west-1. Defaults to us-east-1.
formatcsv or json. Default csv.
delimiterCSV only. One of , ; tab or |. Default comma.
limitRows to return, 1 to 100,000. Default 10,000.

Example

curl https://wapgee.com/api/v1/tools/parquet-to-csv \
  -H "Authorization: Bearer wpg_your_token" \
  -H "Content-Type: application/json" \
  -d '{"url":"s3://my-public-bucket/data/orders.parquet","region":"eu-west-1","limit":100}'

A successful call returns:

{
  "success": true,
  "output": "id,name\n1,Ada\n2,Lin\n",
  "format": "csv",
  "columns": ["id", "name"],
  "rowCount": 2,
  "totalRows": 2,
  "truncated": false
}

What the viewer shows before any export

A Parquet file is not a text file you can paste. Drop it here and the tool reads the footer first: the schema, the row count, and one entry per row group. The schema lists each column's physical type (INT64, DOUBLE, BYTE_ARRAY, and the rest), its logical type when the file has one (UTF8, timestamp, decimal), and whether nulls are allowed. Nested groups are indented under their parent, so a struct or a list is visible before you ask for any cell values.

Row groups are the chunks a query engine reads independently. The table shows how many rows each group holds and how large it is compressed and uncompressed. That is often enough to see whether a file is one giant chunk or already split for scanning.

Large files are not fully decoded for a preview

The preview asks for the first N rows only. The file is read with range requests against the bytes you dropped: the footer for metadata, then the pages that hold those rows. Later row groups stay on disk in the browser and are not decoded until you download a full export. Uncompressed and Snappy files work with the pure JavaScript reader. Gzip, Zstd, and the other codecs are reported from the schema, and the rows are left unread, rather than failing halfway through a decode.

Choose 20, 50, 100, or 500 preview rows. The table and the text box both show that slice. Download full CSV or Download full JSON walks the rest of the file in the browser and saves the result. Nothing is uploaded.

CSV quoting and JSON

CSV export uses the delimiter you pick: comma, semicolon, tab, or pipe. A field is quoted when it contains that delimiter, a double quote, or a line break. Quotes inside a field are doubled, which is the rule spreadsheets already expect. JSON export is an array of objects, pretty-printed, with the same column names as the schema.

A cell
Ada said "hello, world"
CSV field
"Ada said ""hello, world"""

The preview text follows the same rules as the download, limited to the rows on screen. Switch between CSV and JSON without reading the file again.

Nested values, big integers, dates, and binary

Parquet is happy to store things CSV cannot. A struct or a list becomes a JSON value inside the CSV cell, and stays a nested object or array in the JSON download. 64-bit integers are written as decimal strings, not as floating-point numbers, so an id past 2^53 does not get rounded. Dates and timestamps become ISO-8601 strings in UTC. Binary columns become base64, which is reversible and safe to put in either format.

If you started from a spreadsheet and want a Parquet file in the other direction, use CSV to Parquet. For a text-only conversion that never touches Parquet, the CSV and JSON converter is the closer fit.

FAQ

Is my Parquet file uploaded?

No. The file stays in your browser. Metadata, the preview, and the export are all read locally.

Do I have to load the whole file to see a preview?

No. The preview reads the footer and the pages for the first rows you asked for. Later row groups are decoded only when you download a full CSV or JSON export.

Which compression can it read?

Uncompressed and Snappy. Other codecs (Gzip, Zstd, and the rest) are named in the schema, and the rows are not decoded.

How are nested columns and big integers exported?

Nested structs and lists become JSON. Integers that do not fit in a JavaScript number are decimal strings, dates are ISO-8601, and binary columns are base64.