I am doing some computation on the CPU and then I transfer the numbers

Question

0

Asked: May 28, 20262026-05-28T06:49:07+00:00 2026-05-28T06:49:07+00:00

I am doing some computation on the CPU and then I transfer the numbers

0

I am doing some computation on the CPU and then I transfer the numbers to the GPU and do some work there. I want to calculate the total time taken to do the computation on the CPU + the GPU. how do i do so?

Report

Leave an answer
Cancel reply

You must login to add an answer.

Need An Account,

1 Answer

Editorial Team · Answer 1 · 2026-05-28T06:49:07+00:00

When your program starts, in main(), use any system timer to record the time. When your program ends at the bottom of main(), use the same system timer to record the time. Take the difference between time2 and time1. There you go!

There are different system timers you can use, some with higher resolution than others. Rather than discuss those here, I’d suggest you search for “system timer” on the SO site. If you just want any system timer, gettimeofday() works on Linux systems, but it has been superseded by newer, higher-precision functions. As it is, gettimeofday() only measures time in microseconds, which should be sufficient for your needs.

If you can’t get a timer with good enough resolution, consider running your program in a loop many times, timing the execution of the loop, and dividing the measured time by the number of loop iterations.

EDIT:

System timers can be used to measure total application performance, including time used during the GPU calculation. Note that using system timers in this way applies only to real, or wall-clock, time, rather than process time. Measurements based on the wall-clock time must include time spent waiting for GPU operations to complete.

If you want to measure the time taken by a GPU kernel, you have a few options. First, you can use the Compute Visual Profiler to collect a variety of profiling information, and although I’m not sure that it reports time, it must be able to (that’s a basic profiling function). Other profilers – PAPI comes to mind – offer support for CUDA kernels.

Another option is to use CUDA events to record times. Please refer to the CUDA 4.0 Programming Guide where it discusses using CUDA events to measure time.

Yet another option is to use system timers wrapped around GPU kernel invocations. Note that, given the asynchronous nature of kernel invocation returns, you will also need to follow the kernel invocation with a host-side GPU synchronization call such as cudaThreadSynchronize() for this method to be applicable. If you go with this option, I highly recommend calling the kernel in a loop, timing the loop + one synchronization at the end (since synchronization occurs between kernel calls not executing in different streams, cudaThreadSynchronize() is not needed inside the loop), and dividing by the number of iterations.

Sign Up

Sign In

Forgot Password

The Archive Base Latest Questions

I am doing some computation on the CPU and then I transfer the numbers

Leave an answerCancel reply

1 Answer

Leave an answer
Cancel reply