I went to an interview today and was asked this question:
Suppose you have one billion integers which are unsorted in a disk file. How would you determine the largest hundred numbers?
I’m not even sure where I would start on this question. What is the most efficient process to follow to give the correct result? Do I need to go through the disk file a hundred times grabbing the highest number not yet in my list, or is there a better way?
Here’s my initial algorithm:
This has the (very slight) advantage is that there’s no O(n^2) shuffling for the first 100 elements, just an O(n log n) sort and that you very quickly identify and throw away those that are too small. It also uses a binary search (7 comparisons max) to find the correct insertion point rather than 50 (on average) for a simplistic linear search (not that I’m suggesting anyone else proffered such a solution, just that it may impress the interviewer).
You may even get bonus points for suggesting the use of optimised
shiftoperations likememcpyin C provided you can be sure the overlap isn’t a problem.One other possibility you may want to consider is to maintain three lists (of up to 100 integers each):
I’m not sure, but that may end up being more efficient than the continual shuffling.
The merge-sort is a simple selection along the lines of (for merge-sorting lists 1 and 2 into 3):
Simply put, pulling the top 100 values out of the combined list by virtue of the fact they’re already sorted in descending order. I haven’t checked in detail whether that would be more efficient, I’m just offering it as a possibility.
I suspect the interviewers would be impressed with the potential for “out of the box” thinking and the fact that you’d stated that it should be evaluated for performance.
As with most interviews, technical skill is one of the the things they’re looking at.