I have a table that contains maybe 10k to 100k rows and I need varying sets of up to 1 or 2 thousand rows, but often enough a lot less. I want these queries to be as fast as possible and I would like to know which approach is generally smarter:
- Always query for exactly the rows I need with a WHERE clause that’s different all the time.
- Load the whole table into a cache in memory inside my app and search there, syncing the cache regularly
- Always query the whole table (without WHERE clause), let the SQL server handle the cache (it’s always the same query so it can cache the result) and filter the output as needed
I’d like to be agnostic of a specific DB engine for now.
I firmly believe option 1 should be preferred in an initial situation. When you encounter performance problems, you can look on how you could optimize it using caching. (Pre optimization is the root of all evil, Dijkstra once said).
Also, remember that if you would choose option 3, you’ll be sending the complete table-contents over the network as well. This also has an impact on performance .