Credentials
Stored encrypted and decrypted only while a job is running. Use an IAM user scoped to the bucket below.
Temporary credentials
Leave empty for a long-lived IAM user key.
Location
Where browsing starts. Leave it empty to start at the top of the bucket. Saved work addresses objects by their full path, so this is a starting point rather than a restriction.
Remove Duplicate Records from ORC Data
Remove duplicate records from your Apache ORC data quickly
Connectors
Read straight from the systems your data already lives in, instead of uploading a file.
Remove Duplicates
Duplicate rows can cause confusion, errors, and even system failures. This tool scans your Apache ORC file for duplicate entries based on the fields you choose and removes the rows automatically. Whether you're cleaning up customer data, survey responses, or any other dataset, it helps ensure your file is accurate and reliable. You can choose to check for exact duplicates or compare specific columns, giving you full flexibility in how duplicates are identified.
Apache ORC
Apache ORC (Optimized Row Columnar) is a self-describing, columnar file format that supports high compression ratios and fast data retrieval. ORC supports complex types, including structs, lists, maps, and unions. ORC files are divided into blocks of data (stripes) containing statistics (such as min, max, sum, and count) and lightweight indexing which can be used to skip over irrelevant data during queries. ORC also supports predicate pushdown, meaning that filters can be applied as the data is read from disk, reducing the amount of data loaded into memory and processed. Due to its high performance in terms of compression and speed of access, ORC is particularly well-suited for heavy read operations and is commonly used in data warehousing and analytics applications.