Difference between DWH and Datalake?
1) What is Edge Node?
2)What are the client and Cluster-Mode?
3)What is the best approach while running your spark application in Prod?
4)What is Dynamic Memory Allocation in spark and why do we need it?
5)What is shuffle Services and how to enable it?
6)DataFrame and Dataset?
7)Serialization and Deserialization, Encoders and Decoders?
8)Avro, Orc and Parquet differences?
7)Hive ACID properties?
8)What is narrow and Wide Transformation?
9)When stages will be Created in Spark UI?
10)Question from GIT And GITHUB
11)Few Question from Maven
12) What is the process to deploy your code in production?
13)What is Kyro Serialization?
14)What are the optimization techniques used in Spark?
How to connect data node, what is the password type means is that static or dynamic, How you will get the password to connect datanode and apart from that what security you are follow to protect datanode?
2.assume that the data is in JSON format and there is two tables Emp(empid, ename, salary, deptid) & Dept (deptid,dname, loc) Find Dept wise third highest salary in spark?
How will you access bucket.As after performing bucketing on partition table its data will be stored in different directory's file. how will you know which data is in which file of directory .
Need help to understand about Data set Partition in Spark.
How it happens and basis of it considers.
Example -As I have 200mb of data and our block size is 128 mb so basis of what it will distribute data. On some site I saw like we can pass parameter (1 0r 2) to increase number of partition.
Scala>val load=sc.textFile("C:/Desktop/Dept.txt",3)
So in above case partition will be created based on parameter or based on block size.
By default, partitions will be created based on the block size.
Ex: val load=sc.textFile("C:/Desktop/Dept.txt") -
in this case let say, dept.txt file has 200MB -> then will be having two partitions:
128MB one partition and 82 MB as other partition.
you can also exclusively specify the no.of partitions while reading data, to get better performance:
Ex: val load=sc.textFile("C:/Desktop/Dept.txt",3)
in this case let say, dept.txt file has 200MB -> then will be having three equal partitions: around 66.6 MB each partition
==>By default,partitions will be based on "block size".
If we specify the "parameter", Then "partitions" will be the "parameter value" which we specify
i.e blocks will be "shuffled" its in memory(cache) and we get the parttions as per the paramter value.
1)What is your hive engine
2)Why did you chose that particular engine
3) What performance optimizations you did on your data
1=> 2 execution engine mr and spark : set hive.execution.engine=engine. here "engine" is either mr or spark . "mr" is default. and 1more is "tez" is used to increase the speed in petabytes of data.
2=> the answer may be based on specific engine which we have use in our project.
3=> performance optimizations are : Indexing,Fileformats,partitions,bucketing,costbased,vectorization.
1. In a simple scala programe how we find the number of transformations and actions?
2. How we do partition on buckted table??
. How we do partition on buckted table??
I thinkk, we can do bucketing on partition table, but can not do partition on bucketed table.
The use of bucket is to keep fixed number of buckets, to gain performance.
resume display and jobs email and sms
2500
1. what is the difference between map and map partition. explain with example and where to use when?
2. scenario - suppose if an rdd is deleted in spark job how do we can recover it and what is the backend mechanism of recovering rdd's.
3.how do you connect to your cluster using data nodes or edge nodes and what is reason of choosing between the both?
4. have you ever received an error "spaceout" in your datanode?
5. how do you allocate buffer memory to your datanode?
6.how much buffer space have you allocated to your map task and reduce task in your datanode?
7 how do you achieve broadcast join automatically without out doing it manually? and how do you setup your driver program to detect where broadcast join can be good to use and how do you automate the process?
8. how do you acheive in memory caache?
scenario : imagine you are working on cluster and already have cache your rdd and got the output stored in cache now i want to clear the memory space and use that space for caching another rdd? how to achieve this?
9.what are the packages you have worked in scala name the package you have imported in your current project?
10. what modules you have worked in scala and name the module which you have worked till date?
11.Kafka - scenario : suppose your producer is producing more then your consumer can consume , how will you deal such situation and what are your preventive measures to stop data loss?
12. how do you achieve " re-balancing in Kafka and in what way it is use useful?
13. Kafka scenario : suppose producer is writing the data in CSV format and in structure data then how will the consumer will come to know what schema the data is coming in and how to specify and where to specify the schema?
14. how do you manage you offsets?
15 . Kafka : scenario : suppose consumer a has read 10 offsets from the topics and it got failed then how consumer b will pick up offsets and how does it stores the data and what is the mechanism we need configure to achieve this.
16. Hive : Scenario: Imagine we have 2 tables A and B.
B is the master table and A is the table which receives the updates of certain information
so i want to update table B using the latest updated columns based up on the id how do we achieve that and what is the exact query we use?
17. What is use of Row-index and in which scenarios have you used it in hive?
18. what do you know about Ntile.
19. Spark - Scenario : Suppose i m running 10 sql jobs which generally take 10 mins to complete, but one it took 1 hour to complete if this is the case how to you report this error and how will you debug your code and provide a solution for this.
20. what do you about type safety and which frame work has type safety?
21. what are the serializations you have worked on and why do you choose that serialisation explain in detail.
22. how do you achieve performance tuning in spark, apart from use coleasce and reporting, explain me other techniques to achieve performance tuning?
23 . asked me about shell scripting knowledge ,i said i am not much aware of it then the interview got finished after 1.5 hours
What is Rack in HDFS architecture?
Collection of nodes in a hadoop cluster.
2. What is difference between name node and rack or both same?(May Be I missed the class, or sir didn't explain Rack concept)
Obviously both are different.Name node is the master node which maintains the metadata of all the blocks on hdfs and it would be present on any of the racks.where rack is collection of nodes.
2. how many maximum nodes are avalaible in Rack?
3. what is the maximum size of Rack?
40 to 50 data nodes on a rack.. it depends..
4. Can we have more than one replica exist in same rack?
At least one replica is stored on different RACk
Rack: a logical group of the hadoop node which belongs to one location/Area
links
https://training.databricks.com/visualapi.pdf
https://spark.apache.org/docs/latest/rdd-programming-guide.html
https://01205000634306983881.googlegroups.com/attach/9967b2a1daaa5/image.png?part=0.1&view=1&vt=ANaJVrE34iXaQD0lWz-of3agLqfQppPz_wtvPurIGZ4vc4pYXWeQKxtIbYm_HijQGAu1-wD2icKffTcVuvjFq9PIINEoXr0bGtAgov4qt16tVBttM1wy8E4
https://data-flair.training/blogs/hadoop-2-6-multinode-cluster-setup/
1) What is Edge Node?
2)What are the client and Cluster-Mode?
3)What is the best approach while running your spark application in Prod?
4)What is Dynamic Memory Allocation in spark and why do we need it?
5)What is shuffle Services and how to enable it?
6)DataFrame and Dataset?
7)Serialization and Deserialization, Encoders and Decoders?
8)Avro, Orc and Parquet differences?
7)Hive ACID properties?
8)What is narrow and Wide Transformation?
9)When stages will be Created in Spark UI?
10)Question from GIT And GITHUB
11)Few Question from Maven
12) What is the process to deploy your code in production?
13)What is Kyro Serialization?
14)What are the optimization techniques used in Spark?
How to connect data node, what is the password type means is that static or dynamic, How you will get the password to connect datanode and apart from that what security you are follow to protect datanode?
2.assume that the data is in JSON format and there is two tables Emp(empid, ename, salary, deptid) & Dept (deptid,dname, loc) Find Dept wise third highest salary in spark?
How will you access bucket.As after performing bucketing on partition table its data will be stored in different directory's file. how will you know which data is in which file of directory .
Need help to understand about Data set Partition in Spark.
How it happens and basis of it considers.
Example -As I have 200mb of data and our block size is 128 mb so basis of what it will distribute data. On some site I saw like we can pass parameter (1 0r 2) to increase number of partition.
Scala>val load=sc.textFile("C:/Desktop/Dept.txt",3)
So in above case partition will be created based on parameter or based on block size.
By default, partitions will be created based on the block size.
Ex: val load=sc.textFile("C:/Desktop/Dept.txt") -
in this case let say, dept.txt file has 200MB -> then will be having two partitions:
128MB one partition and 82 MB as other partition.
you can also exclusively specify the no.of partitions while reading data, to get better performance:
Ex: val load=sc.textFile("C:/Desktop/Dept.txt",3)
in this case let say, dept.txt file has 200MB -> then will be having three equal partitions: around 66.6 MB each partition
==>By default,partitions will be based on "block size".
If we specify the "parameter", Then "partitions" will be the "parameter value" which we specify
i.e blocks will be "shuffled" its in memory(cache) and we get the parttions as per the paramter value.
1)What is your hive engine
2)Why did you chose that particular engine
3) What performance optimizations you did on your data
1=> 2 execution engine mr and spark : set hive.execution.engine=engine. here "engine" is either mr or spark . "mr" is default. and 1more is "tez" is used to increase the speed in petabytes of data.
2=> the answer may be based on specific engine which we have use in our project.
3=> performance optimizations are : Indexing,Fileformats,partitions,bucketing,costbased,vectorization.
1. In a simple scala programe how we find the number of transformations and actions?
2. How we do partition on buckted table??
. How we do partition on buckted table??
I thinkk, we can do bucketing on partition table, but can not do partition on bucketed table.
The use of bucket is to keep fixed number of buckets, to gain performance.
resume display and jobs email and sms
2500
1. what is the difference between map and map partition. explain with example and where to use when?
2. scenario - suppose if an rdd is deleted in spark job how do we can recover it and what is the backend mechanism of recovering rdd's.
3.how do you connect to your cluster using data nodes or edge nodes and what is reason of choosing between the both?
4. have you ever received an error "spaceout" in your datanode?
5. how do you allocate buffer memory to your datanode?
6.how much buffer space have you allocated to your map task and reduce task in your datanode?
7 how do you achieve broadcast join automatically without out doing it manually? and how do you setup your driver program to detect where broadcast join can be good to use and how do you automate the process?
8. how do you acheive in memory caache?
scenario : imagine you are working on cluster and already have cache your rdd and got the output stored in cache now i want to clear the memory space and use that space for caching another rdd? how to achieve this?
9.what are the packages you have worked in scala name the package you have imported in your current project?
10. what modules you have worked in scala and name the module which you have worked till date?
11.Kafka - scenario : suppose your producer is producing more then your consumer can consume , how will you deal such situation and what are your preventive measures to stop data loss?
12. how do you achieve " re-balancing in Kafka and in what way it is use useful?
13. Kafka scenario : suppose producer is writing the data in CSV format and in structure data then how will the consumer will come to know what schema the data is coming in and how to specify and where to specify the schema?
14. how do you manage you offsets?
15 . Kafka : scenario : suppose consumer a has read 10 offsets from the topics and it got failed then how consumer b will pick up offsets and how does it stores the data and what is the mechanism we need configure to achieve this.
16. Hive : Scenario: Imagine we have 2 tables A and B.
B is the master table and A is the table which receives the updates of certain information
so i want to update table B using the latest updated columns based up on the id how do we achieve that and what is the exact query we use?
17. What is use of Row-index and in which scenarios have you used it in hive?
18. what do you know about Ntile.
19. Spark - Scenario : Suppose i m running 10 sql jobs which generally take 10 mins to complete, but one it took 1 hour to complete if this is the case how to you report this error and how will you debug your code and provide a solution for this.
20. what do you about type safety and which frame work has type safety?
21. what are the serializations you have worked on and why do you choose that serialisation explain in detail.
22. how do you achieve performance tuning in spark, apart from use coleasce and reporting, explain me other techniques to achieve performance tuning?
23 . asked me about shell scripting knowledge ,i said i am not much aware of it then the interview got finished after 1.5 hours
What is Rack in HDFS architecture?
Collection of nodes in a hadoop cluster.
2. What is difference between name node and rack or both same?(May Be I missed the class, or sir didn't explain Rack concept)
Obviously both are different.Name node is the master node which maintains the metadata of all the blocks on hdfs and it would be present on any of the racks.where rack is collection of nodes.
2. how many maximum nodes are avalaible in Rack?
3. what is the maximum size of Rack?
40 to 50 data nodes on a rack.. it depends..
4. Can we have more than one replica exist in same rack?
At least one replica is stored on different RACk
Rack: a logical group of the hadoop node which belongs to one location/Area
links
https://training.databricks.com/visualapi.pdf
https://spark.apache.org/docs/latest/rdd-programming-guide.html
https://01205000634306983881.googlegroups.com/attach/9967b2a1daaa5/image.png?part=0.1&view=1&vt=ANaJVrE34iXaQD0lWz-of3agLqfQppPz_wtvPurIGZ4vc4pYXWeQKxtIbYm_HijQGAu1-wD2icKffTcVuvjFq9PIINEoXr0bGtAgov4qt16tVBttM1wy8E4
https://data-flair.training/blogs/hadoop-2-6-multinode-cluster-setup/
Comments
Post a Comment