Accelerating adaptive inverse distance weighting interpolation algorithm on a graphics processing unit

This paper focuses on designing and implementing parallel adaptive inverse distance weighting (AIDW) interpolation algorithms by using the graphics processing unit (GPU). The AIDW is an improved version of the standard IDW, which can adaptively determine the power parameter according to the data points’ spatial distribution pattern and achieve more accurate predictions than those predicted by IDW. In this paper, we first present two versions of the GPU-accelerated AIDW, i.e. the naive version without profiting from the shared memory and the tiled version taking advantage of the shared memory. We also implement the naive version and the tiled version using two data layouts, structure of arrays and array of aligned structures, on both single and double precision. We then evaluate the performance of parallel AIDW by comparing it with its corresponding serial algorithm on three different machines equipped with the GPUs GT730M, M5000 and K40c. The experimental results indicate that: (i) there is no significant difference in the computational efficiency when different data layouts are employed; (ii) the tiled version is always slightly faster than the naive version; and (iii) on single precision the achieved speed-up can be up to 763 (on the GPU M5000), while on double precision the obtained highest speed-up is 197 (on the GPU K40c). To benefit the community, all source code and testing data related to the presented parallel AIDW algorithm are publicly available.

Download Full-text

Finite element method completely implemented for graphic processor units using parallel algorithm libraries

The International Journal of High Performance Computing Applications ◽

10.1177/1094342017694703 ◽

2017 ◽

Vol 33 (1) ◽

pp. 53-66 ◽

Cited By ~ 1

Author(s):

Franz Pichler ◽

Gundolf Haase

Keyword(s):

Finite Element ◽

Graphics Processing Unit ◽

Computational Cost ◽

Processing Unit ◽

Time Step ◽

Device Architecture ◽

Transient Problems ◽

Speed Up ◽

Automotive Batteries ◽

Graphics Processing

A finite element code is developed in which all of the computationally expensive steps are performed on a graphics processing unit via the THRUST and the PARALUTION libraries. The code focuses on the simulation of transient problems where the repeated computations per time-step create the computational cost. It is used to solve partial and ordinary differential equations as they arise in thermal-runaway simulations of automotive batteries. The speed-up obtained by utilizing the graphics processing unit for every critical step is compared against the single core and the multi-threading solutions which are also supported by the chosen libraries. This way a high total speed-up on the graphics processing unit is achieved without the need for programming a single classical Compute Unified Device Architecture kernel.

Download Full-text

Implementation of a Semi-Implicit Pressure-Based Multigrid Fluid Flow Algorithm on a Graphics Processing Unit

Volume 13: New Developments in Simulation Methods and Software for Engineering Applications; Safety Engineering, Risk Analysis and Reliability Methods; Transportation Systems ◽

10.1115/imece2009-11587 ◽

2009 ◽

Cited By ~ 5

Author(s):

Aaron F. Shinn ◽

S. P. Vanka

Keyword(s):

Stokes Equations ◽

Graphics Processing Unit ◽

Navier Stokes ◽

Processing Unit ◽

Navier Stokes Equations ◽

Driven Cavity ◽

Multigrid Algorithm ◽

Computational Speed ◽

Speed Up ◽

Graphics Processing

A semi-implicit pressure based multigrid algorithm for solving the incompressible Navier-Stokes equations was implemented on a Graphics Processing Unit (GPU) using CUDA (Compute Unified Device Architecture). The multigrid method employed was the Full Approximation Scheme (FAS), which is used for solving nonlinear equations. This algorithm is applied to the 2D driven cavity problem and compared to the CPU version of the code (written in Fortran) to assess computational speed-up.

Download Full-text

Evaluating the Power of GPU Acceleration for IDW Interpolation Algorithm

The Scientific World JOURNAL ◽

10.1155/2014/171574 ◽

2014 ◽

Vol 2014 ◽

pp. 1-8 ◽

Cited By ~ 9

Author(s):

Gang Mei

Keyword(s):

Shared Memory ◽

Inverse Distance Weighting ◽

Experimental Results ◽

Gpu Acceleration ◽

Interpolation Algorithm ◽

Power Parameter ◽

Practical Applications ◽

Distance Weighting ◽

Inverse Distance ◽

Gpu Implementation

We first present two GPU implementations of the standard Inverse Distance Weighting (IDW) interpolation algorithm, the tiled version that takes advantage of shared memory and the CDP version that is implemented using CUDA Dynamic Parallelism (CDP). Then we evaluate the power of GPU acceleration for IDW interpolation algorithm by comparing the performance of CPU implementation with three GPU implementations, that is, the naive version, the tiled version, and the CDP version. Experimental results show that the tilted version has the speedups of 120x and 670x over the CPU version when the power parameterpis set to 2 and 3.0, respectively. In addition, compared to the naive GPU implementation, the tiled version is about two times faster. However, the CDP version is 4.8x∼6.0x slower than the naive GPU version, and therefore does not have any potential advantages in practical applications.

Download Full-text

Parallel computations of the step response of a floor heater with the use of a graphics processing unit. Part 2: results and their evaluation

Bulletin of the Polish Academy of Sciences Technical Sciences ◽

10.2478/bpasts-2013-0102 ◽

2013 ◽

Vol 61 (4) ◽

pp. 949-954 ◽

Cited By ~ 1

Author(s):

J. Gołębiowski ◽

J. Forenc

Keyword(s):

Graphics Processing Unit ◽

Sparse Matrix ◽

Temporal Distribution ◽

Step Response ◽

Processing Unit ◽

Commercial Program ◽

Speed Up ◽

Spatio Temporal ◽

Graphics Processing ◽

Linear Systems Of Equations

Abstract Using models and algorithms presented in the first part of the article, a spatio-temporal distribution of the step response of a floor heater was determined. The results have been presented in the form of heating curves and temperature profiles of the heater in the selected time moments. The computations results were verified through comparing them with the solution obtained with the use of a commercial program - NISA. Additionally, the distribution of the average time constant of thermal processes occurring in the heater was determined. The analysis of the use of a graphics processing unit in numerical computations based on the conjugate gradient method was done. It was proved that the use of a graphics processing unit is profitable in the case of solving linear systems of equations with dense coefficient matrices. In the case of a sparse matrix, the speed-up depends on the number of its non-zero elements.

Download Full-text

Ultrasonic pulse propagation simulation using OpenCL for environment mapping and discovery

The International Journal of High Performance Computing Applications ◽

10.1177/1094342019846290 ◽

2019 ◽

Vol 33 (5) ◽

pp. 1019-1029

Author(s):

Mohammad Y Al-Shorman ◽

Majd M Al-Kofahi

Keyword(s):

Experimental Data ◽

Pulse Propagation ◽

Graphics Processing Unit ◽

Ultrasonic Pulse ◽

Processing Unit ◽

Time Profiles ◽

Simulation Process ◽

Front End ◽

Speed Up ◽

Graphics Processing

A fast, highly parallelized, simulation of unidirectional ultrasonic pulse propagating in a two-dimensional environment is presented. The pulse intensity versus time is recorded using an array of unidirectional ultrasonic receivers located at known locations and arranged in a small circle around the transmitter. To speed up the simulation process, OpenCL 2.0 heterogeneous compute language on a graphics processing unit is used. The simulation result is then compared with experimental data to validate its accuracy. By comparing both simulated and experimental data, the collected intensity–time profiles can be used to map an environment. Environments can be mapped using not only direct reflections but also higher order reflections from objects that are not directly seen by the transmitter. With the help of this simulation, subtle characteristics in an environment, such as a slight tilt or curvature, can be measured. The front end of the simulation is written using C#, while the back end is written using C\C++ and OpenCL.

Download Full-text

Speed up big integer multiplication in the Block Wiedemann on graphics processing unit

Design, Manufacturing and Mechatronics ◽

10.1142/9789813208322_0005 ◽

2017 ◽

Author(s):

Peng-Bo Wu ◽

Jing-Fei Jiang ◽

Yang Zhao

Keyword(s):

Graphics Processing Unit ◽

Processing Unit ◽

Speed Up ◽

Integer Multiplication ◽

Graphics Processing

Download Full-text

Employing graphics processing unit technology, alternating direction implicit method and domain decomposition to speed up the numerical diffusion solver for the biomedical engineering research

International Journal for Numerical Methods in Biomedical Engineering ◽

10.1002/cnm.1444 ◽

2011 ◽

Vol 27 (11) ◽

pp. 1829-1849 ◽

Cited By ~ 11

Author(s):

Beini Jiang ◽

Allan Struthers ◽

Zhe Sun ◽

Zhuo Feng ◽

Xuqian Zhao ◽

...

Keyword(s):

Domain Decomposition ◽

Biomedical Engineering ◽

Graphics Processing Unit ◽

Processing Unit ◽

Alternating Direction Implicit Method ◽

Alternating Direction Implicit ◽

Engineering Research ◽

Alternating Direction ◽

Speed Up ◽

Graphics Processing

Download Full-text

Construction of an optoacoustic image of biological tissues based on an algorithm for a graphics processor

Applied Physics ◽

10.51368/1996-0948-2021-5-106-109 ◽

2021 ◽

pp. 106-109

Author(s):

Denis Kravchuk

Keyword(s):

Gpu Computing ◽

Graphics Processing Unit ◽

Biological Tissues ◽

Ultrasonic Field ◽

Processing Unit ◽

Optoacoustic Imaging ◽

Optoacoustic Interaction ◽

Speed Up ◽

Migration Method ◽

Graphics Processing

The use of optical contrast between different blood particles allows the use of optoacoustic imaging to visualize the distribution of blood particles (erythrocytes, taking into account oxygen saturation), the delivery of drugs to organs through blood vessels. An algorithm for calculating the ultrasonic field obtained as a result of optoacoustic interaction has been developed to speed up calculations on the GPU board. An architecture for fast restoration of an optoacoustic signal based on graphics processing unit (GPU) programming is proposed. The algorithm used in combination with the pre-migration method provides an improvement in the resolution and sharpness of the optoacoustic image of the simulated biological tissues. Thanks to the advanced graphics processing unit (GPU) computing architecture, time-consuming main processing unit (CPU) computing is accelerated with great computational efficiency.

Download Full-text

Scalable Graphics Processing Unit–Based Multiscale Linear Solvers for Reservoir Simulation

SPE Journal ◽

10.2118/203939-pa ◽

2021 ◽

pp. 1-20

Author(s):

A. M. Manea ◽

T. Almani

Keyword(s):

Shared Memory ◽

Reservoir Simulation ◽

Graphics Processing Unit ◽

Parallel Architecture ◽

Multiscale Methods ◽

Massively Parallel ◽

Processing Unit ◽

Multicore Architecture ◽

Graphics Processing ◽

Gpu Architecture

Summary In this work, the scalability of two key multiscale solvers for the pressure equation arising from incompressible flow in heterogeneous porous media, namely, the multiscale finite volume (MSFV) solver, and the restriction-smoothed basis multiscale (MsRSB) solver, are investigated on the graphics processing unit (GPU) massively parallel architecture. The robustness and scalability of both solvers are compared against their corresponding carefully optimized implementation on the shared-memory multicore architecture in a structured problem setting. Although several components in MSFV and MsRSB algorithms are directly parallelizable, their scalability on the GPU architecture depends heavily on the underlying algorithmic details and data-structure design of every step, where one needs to ensure favorable control and data flow on the GPU, while extracting enough parallel work for a massively parallel environment. In addition, the type of algorithm chosen for each step greatly influences the overall robustness of the solver. Thus, we extend the work on the parallel multiscale methods of Manea et al. (2016) to map the MSFV and MsRSB special kernels to the massively parallel GPU architecture. The scalability of our optimized parallel MSFV and MsRSB GPU implementations are demonstrated using highly heterogeneous structured 3D problems derived from the SPE10 Benchmark (Christie and Blunt 2001). Those problems range in size from millions to tens of millions of cells. For both solvers, the multicore implementations are benchmarked on a shared-memory multicore architecture consisting of two packages of Intel® Cascade Lake Xeon Gold 6246 central processing unit (CPU), whereas the GPU implementations are benchmarked on a massively parallel architecture consisting of NVIDIA Volta V100 GPUs. We compare the multicore implementations to the GPU implementations for both the setup and solution stages. Finally, we compare the parallel MsRSB scalability to the scalability of MSFV on the multicore (Manea et al. 2016) and GPU architectures. To the best of our knowledge, this is the first parallel implementation and demonstration of these versatile multiscale solvers on the GPU architecture. NOTE: This paper is published as part of the 2021 SPE Reservoir Simulation Conference Special Issue.

Download Full-text

Analysis of GPU Computation of Parabolic, Bessel, Wright and Riemann Zeta Functions

ITM Web of Conferences ◽

10.1051/itmconf/20214002005 ◽

2021 ◽

Vol 40 ◽

pp. 02005

Author(s):

Ashish A. Jadhav ◽

Abhijeet D. Kalamkar ◽

Pritish A. Gaikwad ◽

Vishwesh Vyawahare ◽

Navin Singhaniya

Keyword(s):

Fractional Calculus ◽

Gpu Computing ◽

Graphics Processing Unit ◽

Zeta Functions ◽

Computation Time ◽

Processing Unit ◽

Mathematical Functions ◽

Speed Up ◽

Sequential Code ◽

Graphics Processing

This paper deals with GPU computing of special mathematical functions that are used in Fractional Calculus. The graphics processing unit (GPU) has grown to be an integral part of nowadays’s mainstream computing structures. The special mathematical functions are an integral part of Fractional Calculus. This paper deals with a novel parallel approach for computing special mathematical functions used in Fractional Calculus. NVIDIA’s GPU hardware is used to speed up the parallel algorithm. A comparison of the sequential code, vectorized code and GPU code is performed. We have successfully reduced the computation time of special mathematical functions using the parallel computing capabilities of GPU.

Download Full-text