Skip to main content

Posts

Showing posts from May, 2019

spark interview QA

Apache Spark is a fast and general-purpose cluster computing system. It provides high-level APIs in Java, Scala, Python and R, and an optimized engine that supports general execution graphs. It also supports a rich set of higher-level tools including Spark SQL for SQL and structured data processing, MLlib for machine learning, GraphX for graph processing, and Spark Streaming."   1)How does Spark relate to Apache Hadoop?   Answer)Spark is a fast and general processing engine compatible with Hadoop data. It can run in Hadoop clusters through YARN or Spark's standalone mode, and it can process data in HDFS, HBase, Cassandra, Hive, and any Hadoop InputFormat. It is designed to perform both batch processing (similar to MapReduce) and new workloads like streaming, interactive queries, and machine learning.  2)Who is using Spark in production?   Answer)As of 2016, surveys show that more than 1000 organizations are using Spark in production. Some of them are listed on the P...

Hive interview QA

Apache Hive data warehouse software facilitates reading, writing, and managing large datasets residing in distributed storage using SQL.  1) What is the definition of Hive? What is the present version of Hive and explain about ACID transactions in Hive?  Answer) Hive is an open source data warehouse system. We can use Hive for analyzing and querying in large data sets of Hadoop files. Its similar to SQL. Hive supports ACID transactions: The full form of ACID is Atomicity, Consistency, Isolation, and Durability. ACID transactions are provided at the row levels, there are Insert, Delete, and Update options so that Hive supports ACID transaction. Insert Delete Update 2)Explain what is a Hive variable. What do we use it for?  Answer)Hive variable is basically created in the Hive environment that is referenced by Hive scripting languages. It provides to pass some values to the hive queries when the query starts executing. It uses the source command....