253 Exam Questions for Professional-Data-Engineer Updated Versions With Test Engine
Pass Professional-Data-Engineer Exam with Updated Professional-Data-Engineer Exam Dumps PDF 2021
Certification Path
The Google Professional Data Engineer Certification is one of the highest level of certification mainly focussing to the professional Data Engineering.
There is no prerequisite for this exam but still it would be best to follow some sequence in order to prove immense knowledge as a Google professional Data Engineer.
You can complete Google Associate Certifications then approach for the professional certification. For more information related to Google cloud certification track Google-certification-path
NEW QUESTION 38
You work for a large real estate firm and are preparing 6 TB of home sales data to be used for machine learning. You will use SQL to transform the data and use BigQuery ML to create a machine learning model. You plan to use the model for predictions against a raw dataset that has not been transformed. How should you set up your workflow in order to prevent skew at prediction time?
- A. Use a BigQuery view to define your preprocessing logic. When creating your model, use the view as your model training data. At prediction time, use BigQuery's ML.EVALUATE clause without specifying any transformations on the raw input data.
- B. When creating your model, use BigQuery's TRANSFORM clause to define preprocessing steps. At prediction time, use BigQuery's ML.EVALUATE clause without specifying any transformations on the raw input data.
- C. Preprocess all data using Dataflow. At prediction time, use BigQuery's ML.EVALUATE clause without specifying any further transformations on the input data.
- D. When creating your model, use BigQuery's TRANSFORM clause to define preprocessing steps. Before requesting predictions, use a saved query to transform your raw input data, and then use ML.EVALUATE.
Answer: D
NEW QUESTION 39
Flowlogistic wants to use Google BigQuery as their primary analysis system, but they still have Apache Hadoop and Spark workloads that they cannot move to BigQuery. Flowlogistic does not know how to store the data that is common to both workloads. What should they do?
- A. Store the common data encoded as Avro in Google Cloud Storage.
- B. Store the common data in BigQuery as partitioned tables.
- C. Store he common data in the HDFS storage for a Google Cloud Dataproc cluster.
- D. Store the common data in BigQuery and expose authorized views.
Answer: D
NEW QUESTION 40
Cloud Bigtable is a recommended option for storing very large amounts of
____________________________?
- A. multi-keyed data with very low latency
- B. single-keyed data with very high latency
- C. single-keyed data with very low latency
- D. multi-keyed data with very high latency
Answer: C
Explanation:
Cloud Bigtable is a sparsely populated table that can scale to billions of rows and thousands of columns, allowing you to store terabytes or even petabytes of data. A single value in each row is indexed; this value is known as the row key. Cloud Bigtable is ideal for storing very large amounts of single-keyed data with very low latency. It supports high read and write throughput at low latency, and it is an ideal data source for MapReduce operations.
Reference: https://cloud.google.com/bigtable/docs/overview
NEW QUESTION 41
The Dataflow SDKs have been recently transitioned into which Apache service?
- A. Apache Kafka
- B. Apache Hadoop
- C. Apache Beam
- D. Apache Spark
Answer: C
Explanation:
Dataflow SDKs are being transitioned to Apache Beam, as per the latest Google directive
https://cloud.google.com/dataflow/docs/
NEW QUESTION 42
You are designing storage for very large text files for a data pipeline on Google Cloud. You want to support ANSI SQL queries. You also want to support compression and parallel load from the input locations using Google recommended practices. What should you do?
- A. Compress text files to gzip using the Grid Computing Tools. Use BigQuery for storage and query.
- B. Transform text files to compressed Avro using Cloud Dataflow. Use BigQuery for storage and query.
- C. Compress text files to gzip using the Grid Computing Tools. Use Cloud Storage, and then import into Cloud Bigtable for query.
- D. Transform text files to compressed Avro using Cloud Dataflow. Use Cloud Storage and BigQuery permanent linked tables for query.
Answer: C
NEW QUESTION 43
You have a query that filters a BigQuery table using a WHERE clause on timestamp and ID columns. By using bq query - -dry_run you learn that the query triggers a full scan of the table, even though the filter on timestamp and ID select a tiny fraction of the overall data. You want to reduce the amount of data scanned by BigQuery with minimal changes to existing SQL queries. What should you do?
- A. Recreate the table with a partitioning column and clustering column.
- B. Create a separate table for each ID.
- C. Use the LIMIT keyword to reduce the number of rows returned.
- D. Use the bq query - -maximum_bytes_billed flag to restrict the number of bytes billed.
Answer: C
NEW QUESTION 44
You work for an advertising company, and you've developed a Spark ML model to predict click-through rates at advertisement blocks. You've been developing everything at your on-premises data center, and now your company is migrating to Google Cloud. Your data center will be closing soon, so a rapid lift-and-shift migration is necessary. However, the data you've been using will be migrated to migrated to BigQuery. You periodically retrain your Spark ML models, so you need to migrate existing training pipelines to Google Cloud. What should you do?
- A. Use Cloud ML Engine for training existing Spark ML models
- B. Use Cloud Dataproc for training existing Spark ML models, but start reading data directly from BigQuery
- C. Rewrite your models on TensorFlow, and start using Cloud ML Engine
- D. Spin up a Spark cluster on Compute Engine, and train Spark ML models on the data exported from BigQuery
Answer: A
Explanation:
Explanation
NEW QUESTION 45
Case Study: 2 - MJTelco
Company Overview
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost. Their management and operations teams are situated all around the globe creating many-to- many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments ?development/test, staging, and production ?
to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements
Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community. Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
Provide reliable and timely access to data for analysis from distributed research workers Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements
Ensure secure and efficient transport and storage of telemetry data Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately
100m records/day
Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement
The project is too large for us to maintain the hardware and software required for the data and analysis.
Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
Given the record streams MJTelco is interested in ingesting per day, they are concerned about the cost of Google BigQuery increasing. MJTelco asks you to provide a design solution. They require a single large data table called tracking_table. Additionally, they want to minimize the cost of daily queries while performing fine-grained analysis of each day's events. They also want to use streaming ingestion. What should you do?
- A. Create sharded tables for each day following the pattern tracking_table_YYYYMMDD.
- B. Create a table called tracking_table with a TIMESTAMP column to represent the day.
- C. Create a partitioned table called tracking_table and include a TIMESTAMP column.
- D. Create a table called tracking_table and include a DATE column.
Answer: C
NEW QUESTION 46
Which of the following statements about Legacy SQL and Standard SQL is not true?
- A. Standard SQL is the preferred query language for BigQuery.
- B. One difference between the two query languages is how you specify fully-qualified table names (i.e.
table names that include their associated project name). - C. If you write a query in Legacy SQL, it might generate an error if you try to run it with Standard SQL.
- D. You need to set a query language for each dataset and the default is Standard SQL.
Answer: D
Explanation:
Explanation
You do not set a query language for each dataset. It is set each time you run a query and the default query language is Legacy SQL.
Standard SQL has been the preferred query language since BigQuery 2.0 was released.
In legacy SQL, to query a table with a project-qualified name, you use a colon, :, as a separator. In standard SQL, you use a period, ., instead.
Due to the differences in syntax between the two query languages (such as with project-qualified table names), if you write a query in Legacy SQL, it might generate an error if you try to run it with Standard SQL.
Reference:
https://cloud.google.com/bigquery/docs/reference/standard-sql/migrating-from-legacy-sql
NEW QUESTION 47
Your organization has been collecting and analyzing data in Google BigQuery for 6 months. The majority of the data analyzed is placed in a time-partitioned table named events_partitioned. To reduce the cost of queries, your organization created a view called events, which queries only the last 14 days of data. The view is described in legacy SQL. Next month, existing applications will be connecting to BigQuery to read the eventsdata via an ODBC connection. You need to ensure the applications can connect. Which two actions should you take? (Choose two.)
- A. Create a new view over events using standard SQL
- B. Create a service account for the ODBC connection to use for authentication
- C. Create a new partitioned table using a standard SQL query
- D. Create a new view over events_partitioned using standard SQL
- E. Create a Google Cloud Identity and Access Management (Cloud IAM) role for the ODBC connection and shared "events"
Answer: A,E
NEW QUESTION 48
Your company produces 20,000 files every hour. Each data file is formatted as a comma separated values
(CSV) file that is less than 4 KB. All files must be ingested on Google Cloud Platform before they can be
processed. Your company site has a 200 ms latency to Google Cloud, and your Internet connection
bandwidth is limited as 50 Mbps. You currently deploy a secure FTP (SFTP) server on a virtual machine in
Google Compute Engine as the data ingestion point. A local SFTP client runs on a dedicated machine to
transmit the CSV files as is. The goal is to make reports with data from the previous day available to the
executives by 10:00 a.m. each day. This design is barely able to keep up with the current volume, even
though the bandwidth utilization is rather low.
You are told that due to seasonality, your company expects the number of files to double for the next three
months. Which two actions should you take? (Choose two.)
- A. Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer
Service to transfer on-premices data to the designated storage bucket. - B. Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble
the CSV files in the cloud upon receiving them. - C. Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in
parallel. - D. Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.
- E. Introduce data compression for each file to increase the rate file of file transfer.
Answer: A,C
NEW QUESTION 49
Which of the following statements about Legacy SQL and Standard SQL is not true?
- A. Standard SQL is the preferred query language for BigQuery.
- B. One difference between the two query languages is how you specify fully-qualified table names (i.e.
table names that include their associated project name). - C. If you write a query in Legacy SQL, it might generate an error if you try to run it with Standard SQL.
- D. You need to set a query language for each dataset and the default is Standard SQL.
Answer: D
Explanation:
You do not set a query language for each dataset. It is set each time you run a query and the default query language is Legacy SQL.
Standard SQL has been the preferred query language since BigQuery 2.0 was released. In legacy SQL, to query a table with a project-qualified name, you use a colon, :, as a separator. In standard SQL, you use a period, ., instead.
Due to the differences in syntax between the two query languages (such as with project-qualified table names), if you write a query in Legacy SQL, it might generate an error if you try to run it with Standard SQL.
Reference:
https://cloud.google.com/bigquery/docs/reference/standard-sql/migrating-from-legacy-sql
NEW QUESTION 50
Flowlogistic Case Study
Company Overview
Flowlogistic is a leading logistics and supply chain provider. They help businesses throughout the world manage their resources and transport them to their final destination. The company has grown rapidly, expanding their offerings to include rail, truck, aircraft, and oceanic shipping.
Company Background
The company started as a regional trucking company, and then expanded into other logistics market.
Because they have not updated their infrastructure, managing and tracking orders and shipments has become a bottleneck. To improve operations, Flowlogistic developed proprietary technology for tracking shipments in real time at the parcel level. However, they are unable to deploy it because their technology stack, based on Apache Kafka, cannot support the processing volume. In addition, Flowlogistic wants to further analyze their orders and shipments to determine how best to deploy their resources.
Solution Concept
Flowlogistic wants to implement two concepts using the cloud:
Use their proprietary technology in a real-time inventory-tracking system that indicates the location of
their loads
Perform analytics on all their orders and shipment logs, which contain both structured and unstructured
data, to determine how best to deploy resources, which markets to expand info. They also want to use predictive analytics to learn earlier when a shipment will be delayed.
Existing Technical Environment
Flowlogistic architecture resides in a single data center:
Databases
8 physical servers in 2 clusters
- SQL Server - user data, inventory, static data
3 physical servers
- Cassandra - metadata, tracking messages
10 Kafka servers - tracking message aggregation and batch insert
Application servers - customer front end, middleware for order/customs
60 virtual machines across 20 physical servers
- Tomcat - Java services
- Nginx - static content
- Batch servers
Storage appliances
- iSCSI for virtual machine (VM) hosts
- Fibre Channel storage area network (FC SAN) - SQL server storage
- Network-attached storage (NAS) image storage, logs, backups
Apache Hadoop /Spark servers
- Core Data Lake
- Data analysis workloads
20 miscellaneous servers
- Jenkins, monitoring, bastion hosts,
Business Requirements
Build a reliable and reproducible environment with scaled panty of production.
Aggregate data in a centralized Data Lake for analysis
Use historical data to perform predictive analytics on future shipments
Accurately track every shipment worldwide using proprietary technology
Improve business agility and speed of innovation through rapid provisioning of new resources
Analyze and optimize architecture for performance in the cloud
Migrate fully to the cloud if all other requirements are met
Technical Requirements
Handle both streaming and batch data
Migrate existing Hadoop workloads
Ensure architecture is scalable and elastic to meet the changing demands of the company.
Use managed services whenever possible
Encrypt data flight and at rest
Connect a VPN between the production data center and cloud environment
SEO Statement
We have grown so quickly that our inability to upgrade our infrastructure is really hampering further growth and efficiency. We are efficient at moving shipments around the world, but we are inefficient at moving data around.
We need to organize our information so we can more easily understand where our customers are and what they are shipping.
CTO Statement
IT has never been a priority for us, so as our data has grown, we have not invested enough in our technology. I have a good staff to manage IT, but they are so busy managing our infrastructure that I cannot get them to do the things that really matter, such as organizing our data, building the analytics, and figuring out how to implement the CFO' s tracking technology.
CFO Statement
Part of our competitive advantage is that we penalize ourselves for late shipments and deliveries. Knowing where out shipments are at all times has a direct correlation to our bottom line and profitability.
Additionally, I don't want to commit capital to building out a server environment.
Flowlogistic's CEO wants to gain rapid insight into their customer base so his sales team can be better informed in the field. This team is not very technical, so they've purchased a visualization tool to simplify the creation of BigQuery reports. However, they've been overwhelmed by all the data in the table, and are spending a lot of money on queries trying to find the data they need. You want to solve their problem in the most cost-effective way. What should you do?
- A. Export the data into a Google Sheet for virtualization.
- B. Create identity and access management (IAM) roles on the appropriate columns, so only they appear in a query.
- C. Create a view on the table to present to the virtualization tool.
- D. Create an additional table with only the necessary columns.
Answer: C
NEW QUESTION 51
You work for a large fast food restaurant chain with over 400,000 employees. You store employee information in Google BigQuery in a Users table consisting of a FirstName field and a LastName field. A member of IT is building an application and asks you to modify the schema and data in BigQuery so the application can query a FullName field consisting of the value of the FirstName field concatenated with a space, followed by the value of the LastName field for each employee. How can you make that data available while minimizing cost?
- A. Add a new column called FullName to the Users table. Run an UPDATE statement that updates the FullName column for each user with the concatenation of the FirstName and LastName values.
- B. Create a view in BigQuery that concatenates the FirstName and LastName field values to produce the FullName.
- C. Create a Google Cloud Dataflow job that queries BigQuery for the entire Users table, concatenates the FirstName value and LastName value for each user, and loads the proper values for FirstName, LastName, and FullName into a new table in BigQuery.
- D. Use BigQuery to export the data for the table to a CSV file. Create a Google Cloud Dataproc job to process the CSV file and output a new CSV file containing the proper values for FirstName, LastName and FullName. Run a BigQuery load job to load the new CSV file into BigQuery.
Answer: D
Explanation:
Import and Export to Bigquery from Cloud Storage is FREE. Also, when u store the csv files, Cloud Storage is cheaper than Bigquery. For processing Dataproc is cheaper than Dataflow.
NEW QUESTION 52
The marketing team at your organization provides regular updates of a segment of your customer dataset.
The marketing team has given you a CSV with 1 million records that must be updated in BigQuery. When you use the UPDATE statement in BigQuery, you receive a quotaExceeded error. What should you do?
- A. Increase the BigQuery UPDATE DML statement limit in the Quota management section of the Google Cloud Platform Console.
- B. Reduce the number of records updated each day to stay within the BigQuery UPDATE DML statement limit.
- C. Import the new records from the CSV file into a new BigQuery table. Create a BigQuery job that merges the new records with the existing records and writes the results to a new BigQuery table.
- D. Split the source CSV file into smaller CSV files in Cloud Storage to reduce the number of BigQuery UPDATE DML statements per BigQuery job.
Answer: C
Explanation:
https://cloud.google.com/blog/products/gcp/performing-large-scale-mutations-in-bigquery
NEW QUESTION 53
You are working on a sensitive project involving private user dat
- A. Grant the consultant the Cloud Dataflow Developer role on the project.
- B. You have set up a project on Google Cloud Platform to house your work internally. An external consultant is going to assist with coding a complex transformation in a Google Cloud Dataflow pipeline for your project. How should you maintain users' privacy?
- C. Grant the consultant the Viewer role on the project.
- D. Create an anonymized sample of the data for the consultant to work with in a different project.
- E. Create a service account and allow the consultant to log on with it.
Answer: A
NEW QUESTION 54
......
Operationalize ML Models
- Select the Relevant Training & Service Infrastructure: The consideration for this topic includes distributed versus single machine, hardware accelerators (such as TPU and GPU), and edge compute usage;
- Leverage Pre-Built Machine Learning Models as a Service: It covers one’s knowledge and skills in customizing machine learning APIs, including Auto ML text and Auto ML Vision. It also covers the conversational experiences, such as Dialogflow as well as machine learning APIs, including Speech API and Vision API;
- Deploy Machine Learning Pipelines: This objective requires your competence in ingesting relevant data, continuous evaluation, and retraining of ML models (Kuberflow, BigQuery Machine Learning, Cloud Machine Learning Engine, and Spark Machine Learning);
- Measure, Troubleshoot & Monitor Machine Learning Models: The focus of this subtopic includes the effect of dependencies on machine learning models. It will also measure the examinees’ understanding of machine learning terminologies, such as features, regression, labels, classification, models, recommendation, evaluation metrics, and unsupervised & supervised learning. Moreover, it will also assess their knowledge of common sources of error such as assumptions regarding data.
Professional-Data-Engineer Exam Dumps - Free Demo & 365 Day Updates: https://www.validtorrent.com/Professional-Data-Engineer-valid-exam-torrent.html
Free Sales Ending Soon - Use Real Professional-Data-Engineer PDF Questions: https://drive.google.com/open?id=1FGnmyP8U7HFbuQX47RE9uh63H4RjyhW3