Marcus Warnerfjord
12/09/2025, 11:43 AMRavi Kumar Pilla
12/09/2025, 4:33 PMWHERE prediction IS NULL
pandas.SQLTableDataset Use save_args: {if_exists: "append"} for new rows
If you want to get support for incremental checkpoints, you can explore Incremental Datasets or Partitioned Datasets
If you want incremental native storage, you can explore
DuckDB - ibis.TableDataset - Local/lightweight, great pandas interop
Delta Lake - spark.DeltaTableDataset - Distributed, ACID transactions, time travel
Polars - polars.LazyPolarsDataset - Memory-efficient, fast local analytics
I am not sure if this is helpful for your use-case but we have a native BioSequenceDataset
I will also let the community provide their suggestions. Please let me know if you need further information. Thank youMarcus Warnerfjord
12/12/2025, 9:10 AM