Wapgee Logowapgee

CSV to Parquet

API

Paste or drop a CSV and download a Parquet file. Types are inferred, nulls stay optional, and you can override a column before building. Nothing is uploaded.

CSV170 B
Schema3 rows
ColumnTypeNulls
idrequired
namerequired
qtyoptional
pricerequired
activerequired
createdrequired

Changing a type clears the built file. Build again, then download. Nothing is uploaded.

API

Call this tool over HTTP with a personal API token. The tool above keeps running in your browser. The API is a separate, server-side path. See API documentation for tokens, authentication, limits, and errors shared by every endpoint.

Convert CSV text, or a CSV at a public link, into a Parquet file. Types are inferred the same way as on the tool page.

POST /api/v1/tools/csv-to-parquet

Daily allowance: 10 successful requests, shared by every token on your account. Resets at midnight UTC.

Send exactly one of input (CSV text, up to 10 MB) or url (a public http(s) URL or s3://bucket/key, up to 25 MB). A successful response is the Parquet file itself, with Content-Type application/vnd.apache.parquet. Errors are JSON. Invalid CSV returns 400, and input over the limit returns 413.

Request body

FieldDescription
inputThe CSV text. Use this or url.
urlA public http(s) URL or s3://bucket/key address of the CSV. Use this or input.
regionAWS region for an s3:// address. Only with url.
headerWhether the first row is a header. Default true.
delimiterauto, comma, semicolon, tab or pipe. Default auto.
compressionSNAPPY or UNCOMPRESSED. Default SNAPPY.
columnTypesObject mapping a column name to string, int, double, boolean or timestamp, to override inference.

Example

curl https://wapgee.com/api/v1/tools/csv-to-parquet \
  -H "Authorization: Bearer wpg_your_token" \
  -H "Content-Type: application/json" \
  -o data.parquet \
  -d '{"input":"id,name\n1,Ada\n2,Lin","compression":"SNAPPY"}'

A successful call returns:

HTTP/1.1 200 OK
Content-Type: application/vnd.apache.parquet
Content-Disposition: attachment; filename="data.parquet"

(binary Parquet file)

Why convert a CSV to Parquet

CSV is what spreadsheets and exports hand you: one record per line, everything stored as text. Parquet is what warehouses, Spark, DuckDB, and most data lakes want instead. Columns are typed, repeated values compress well, and a reader can fetch one column without scanning the whole file. This tool does that conversion in your browser. The CSV is parsed here, the Parquet file is built here, and the download never goes through a server.

Paste a table, drop a .csv file, or start from the example. Set the delimiter if auto-detect guesses wrong (semicolons from European Excel are the usual case), and turn the header row off when the file does not have one. Column names then become column_1, column_2, and so on.

How column types are chosen

CSV has no types, so each column is inferred from its cells. A column of true and false becomes boolean. Whole numbers become int. Numbers with a decimal or an exponent become double, and a mix of whole numbers and decimals widens to double. Values like 2024-01-02 or 2024-01-02T03:04:05Z become timestamps. Anything else, including zip codes with a leading zero, stays a string. If the cells disagree (a number in one row and a word in another), the column stays a string so nothing is silently corrupted.

Empty cells, and the token null, are written as null. One null is enough to mark the column optional. A column with a value in every row is required. The schema table shows that choice and lets you override the type before you build. If a cell cannot be read as the type you picked, the build stops and names the column, the row, and the value.

CSV
zip,qty,price,active
02115,3,12.5,true
10001,,2,false
Inferred schema
zip     string     required
qty     int        optional
price   double     required
active  boolean    required

Integers past 2^53

JavaScript numbers are floating point. They stop counting by ones past 9,007,199,254,740,991 (2^53). A 64-bit id or a timestamp in microseconds will not survive a round trip through an ordinary number. Int columns are written as Parquet INT64, and any value outside that safe range is kept as a BigInt, so the file still holds the exact integer. The schema note calls out which column needed it.

Timestamps without a timezone are read as UTC, so the same CSV produces the same instants in every browser. They are stored as milliseconds. Compression is Snappy by default, which is the codec most Parquet readers expect, or none if you want a file you can inspect byte for byte. Both are pure JavaScript. Nothing here loads WebAssembly.

Size, quoting, and what leaves the browser

Files over about 50 MB are refused, so a huge drop does not freeze the tab. Above 1 MB the schema waits until you click Build, for the same reason. Quoted fields are handled by PapaParse: a comma, a quote, or a newline inside a cell stays inside that cell. Duplicate header names are renamed (name, name_2), and a short row is padded with null rather than shifted into the wrong column.

The download button shows the Parquet file size once the file has been built. Changing a type, the delimiter, or the CSV itself clears that file so you cannot download a schema you no longer see. To open the result, use the Parquet to CSV tool, or load it in DuckDB, pandas, or Spark. For a JSON copy of the same table, the CSV and JSON converter stays in text.

FAQ

Is my CSV uploaded?

No. Parsing and the Parquet write both run in your browser. The file is downloaded from this page and never sent to a server.

What is the maximum file size?

About 50 MB. Larger files are refused so the tab does not freeze. Above 1 MB the schema is inferred when you click Build rather than on every keystroke.

Which compression is used?

Snappy by default, or no compression if you pick None. Both are written in pure JavaScript. Snappy is what most Parquet readers expect.

Why did a long id stay exact?

Int columns are 64-bit. Values past 2^53, which JavaScript numbers cannot represent, are stored as BigInt so the Parquet file keeps the original integer.