Browse all practice questions for the CertNexus Certified Data Science Practitioner (CDSP) Practice Exam. Search by topic, open any question and review its full explanation, then test yourself in the practice quiz.

CertNexus Certified Data Science Practitioner (CDSP) Practice Exam course image
More practice questions

These questions are part of the practice quiz. Start practicing

  • Which cost function calculates the average difference between estimated and actual values without factoring in their signs?
  • What does elastic net regression combine in its approach?
  • Which statistical test is used to compare the means of two distributions when the population standard deviation is known?
  • What is the splitting metric used in decision trees that assesses the purity of nodes?
  • What is the role of a threshold in a binary classification model?
  • What type of error occurs when the chosen model is overly complex and captures noise instead of the underlying trend?
  • What is a key characteristic of leptokurtic distributions?
  • Which type of analysis involves summarizing patterns and relationships in data using statistical measures and visualizations?
  • What is the measure of how often the positive identifications made by a learning model are true positives?
  • What term describes a model that performs well on any new datasets it might encounter?
  • In an experiment, what is the term for a variable that can affect the dependent variable?
  • Which model only considers a single variable in its prediction?
  • Which concept refers to the trend that, as more data is added to a model, the model's performance reaches an optimal point beyond which additional data has a negligible effect?
  • What hyperparameter determines how deep a decision tree can grow?
  • What term describes incorrect or missing values in a dataset?
  • In data science, what is often the goal of using a stochastic model?
  • What is a primary function of ridge regression?
  • Which type of data holds number values that express magnitude?
  • In the context of machine learning, what does the term "k" in k-fold cross-validation refer to?
  • What correction is applied when performing variance calculations on a sample by subtracting 1 from the total number of values?
  • What method would primarily be used for reducing the dimensionality of a categorical dataset?
  • What property of a dataset is displayed when there is a high density of values clustered at one end of the distribution?
  • Which statistical concept relates to determining how many standard deviations a value is from the mean?
  • Which type of algorithms are characterized by generating a potentially infinite number of model parameters?
  • What is the primary characteristic of a model that suffers from overfitting?
  • Which of the following lists includes different types of data storage solutions?
  • What is the term for the calculation involving the average, mode, and standard deviation that indicates skewness?
  • What is the term for a sequential set of processing that automates the data science process by feeding the output of one process into the input of the next process?
  • What is the term for the minimum number of samples required to be a leaf node in a decision tree?
  • What does the mean squared error (MSE) function primarily measure in a machine learning model?
  • What are hyperparameters in the context of machine learning?
  • What aspect of model evaluation is concerned with the proportion of actual positives correctly identified?
  • What is the process of making predictions about future events based on past event analysis called?
  • What type of graphical representation shows data points in relation to their geographical location?
  • What is the relation between ARIMA and time series analysis?
  • Which term refers to the practice of giving different weights or importance to different components within a model?
  • Which statistical test compares the effects of categorical variables?
  • What is the key characteristic of Ridge regularization?
  • What is the purpose of the Area Under ROC Curve (AUC) metric?
  • Which term refers to the measure of variability that captures the range of the middle half of data values?
  • What regularization method uses the l2 norm for its regularization term?
  • What is represented on a lift chart?
  • Which of the following best describes continuous variables?
  • Which hyperparameter tuning method randomly selects combinations of hyperparameters?
  • What is an example of data that is not classified as big data?
  • Which of the following describes project deliverables in a project scope?
  • What is the main objective of dimensionality reduction in data science?
  • What type of data consists of numerical values that stand for magnitude?
  • What is a characteristic of multi-label classification?
  • What type of data representation involves ordering observations according to changes over time?
  • Which term refers to the transformation and loading process of data into a destination?
  • What term describes a mathematical relationship between two variables?
  • What does HIPAA stand for in relation to healthcare regulations?
  • What is the measure of how many positive instances a model identifies compared to all relevant instances called?
  • What U.S. law was enacted in 1996 to regulate healthcare practices?
  • What is an area plot in data visualization?
  • What does the null hypothesis assume in statistical testing?
  • What is the process of combining and preparing data from multiple sources called?
  • What is a regular expression used for?
  • What does AUC stand for in the context of model evaluation?
  • What issue arises when a model is too simplistic, resulting in an inability to derive relevant insights from new data?
  • Which statistical test is used to compare the means of two distributions when the population standard deviation is unknown?
  • What is the term for bias introduced when the training dataset is not representative of the target population?
  • What does RMSE stand for in the context of evaluating model performance?
  • What does the term 'dimensions' refer to in the context of a model?
  • What is the term used for the phenomenon where a machine learning model's performance deteriorates over time due to changes in the patterns of data?
  • What is the technique called that improves a model’s ability to make estimations by generating features?
  • What does the term "overfitting" refer to in machine learning?
  • What statistical test is used to compare the means of multiple distributions?
  • What is the term for a distribution that has more than one peak?
  • What type of regression analysis provides a classification probability between 0 and 1?
  • Data binning helps in managing which aspect of the dataset?
  • Which function refers to how independent variables relate to the dependent variables to best meet expectations?
  • What does a z-score represent in statistics?
  • Which of the following is NOT a feature of ANOVA?
  • What optimization method uses past samples to influence where future sampling occurs in order to find the next optimal sample space?
  • What type of regression analysis deals with an independent and a dependent variable in a linear relationship?
  • What type of graph is used to illustrate the relationship between two quantitative variables?
  • What kind of learning uses data that is difficult to search, filter, or extract?
  • What is PCI DSS best known for?
  • Which type of kurtosis is characterized by a narrow peak and heavy tails?
  • What process involves placing the values of continuous variables into specific, discrete intervals?
  • What type of plot is used to show the distribution of a numerical value through probability density?
  • Which of the following refers to a diagram that represents a tree-like hierarchy?
  • What is the process called that simplifies a dataset by removing redundant or irrelevant features?
  • Which of the following best describes supervised learning?
  • What does the coefficient of determination (R^2) indicate?
  • What does feature engineering primarily aim to enhance in a machine learning model?
  • Which type of join returns all records from the first dataset and only the matching records from the second dataset?
  • What is the primary goal of tokenization in text processing?
  • Which type of join returns all records from both datasets, matching where possible?
  • What approach would you use to estimate missing values in a dataset?
  • What is the term for the process of cleaning and organizing raw data into a usable format?
  • What tool is used to visualize the results of a classification problem?
  • Which of the following best describes platykurtic distributions?
  • What k-fold cross-validation method uses all data points in the dataset as folds?
  • In machine learning, what does "iterative" refer to in the context of gradient descent?
  • Which testing method allows researchers to determine if there are statistical differences between group means?
  • What classification problem involves assigning multiple labels to a single data example?
  • What type of plot is characterized by connecting data points in order with a series of lines?
  • What is the primary purpose of the Gini index in decision trees?
  • Which of the following visualizations is used to show central tendency and variation in data distribution?
  • Which term describes the difference between the smallest and largest values in a dataset?
  • What is a confidence interval?
  • What do attributes (or features) contain in a model?
  • What is the technique of condensing a language vocabulary into smaller dimensional vectors called?
  • What aspect of a sample set does the sample mean represent?
  • What does the area under the ROC curve represent?
  • Which term refers to the probability distribution that has a bell-shaped curve?
  • Which type of data expresses categories without a meaningful order?
  • What is the primary purpose of data visualizations?
  • Which type of variable is characterized by having countable, limited values and finite gaps between them?
  • What type of classification approach uses support vector machines to maximize the margin distance?
  • What is used to assess the skill and performance of a machine learning model?
  • What technique samples the training dataset for each individual tree while allowing data examples to appear in multiple models?
  • What term is used to describe a model that is deemed useful for its intended task?
  • What metric is often used to assess the performance of classification models alongside F1 score?
  • Which method visually compares the change in a model's performance against the number of data examples used?
  • What is the measure that represents the square root of variance?
  • What term describes the merging of development and operations in a tech context?
  • Which term describes parameters that are typically set before the training of a machine learning model begins?
  • What term refers to the extent to which data varies across all values in a dataset?
  • What is the term for the process of closely examining data to reveal new insights?
  • What term is defined as the average of all numbers in a data set?
  • What is the term for variables that change indirectly in an experiment?
  • When a model cannot capture the underlying trends in the data, it is said to be:
  • What process involves adjusting hyperparameters used by an algorithm to improve model performance?
  • What law states that when a measure becomes a target, it ceases to be a good measure?
  • What is the main purpose of latent class analysis?
  • Which of the following terms best describes the diversity of outcomes that a model can produce due to variability in the data?
  • What is the role of the cost function in machine learning?
  • What is the purpose of stratified k-fold cross-validation?
  • What type of distribution illustrates the frequency of outcomes for a specific random variable sample?
  • Which metric provides the weighted average of precision and recall?
  • Which of the following best describes 'regularization'?
  • Which technique is used to assess the performance of a classification model?
  • What statistical measure provides an indication of how closely the data points cluster around the mean?
  • What analytical method assesses how well a data point fits within a cluster relative to others?
  • Which plot visually represents the range measurements such as Median, Q1, Q3, Minimum, and Maximum?
  • What is the name of the cross-validation method that splits a dataset into training and test sets?
  • What type of kurtosis has a value equal to 3?
  • What is the formula used to calculate kurtosis?
  • In data science, what is the significance of 'data preprocessing'?
  • In project management, what term describes a detailed outline of all aspects of a project, including constraints?
  • What is the term for a decision boundary in support vector machines that has parallel and equidistant lines on either side?
  • What term refers to a variable that a data science practitioner seeks to learn more about?
  • What is a function that represents the distribution of a random variable as a symmetrical bell-shaped graph?
  • What cross-validation method is defined by leaving one participant out to minimize performance issues?
  • Which technique involves scaling features so that the lowest value is 0 and the highest is 1?
  • In data analysis, what term describes the use of numerical values to summarize data patterns?
  • Which of the following is NOT a regularization technique?
  • What system uses k-nearest neighbor for classification of data examples?
  • What bias arises when training data excludes participants who have dropped out over time?
  • Data that holds categorical values is referred to as what type of data?
  • What process involves taking data as input and representing it in a certain structure or syntax?
  • In gradient descent, what is referred to as the learning rate?
  • Which analysis method is primarily focused on predicting future outcomes based on current data?
  • Which set of statistical parameters is used to measure a distribution?
  • Gradient boosting is primarily used for which type of modeling?
  • Which measure reflects the separation between clusters in a dataset?
  • What type of test compares two different values of the same variable to determine the most effective one?
  • What characteristic defines a unimodal distribution?
  • What is a sample set?
  • How is variance calculated for a sample set?
  • Which algorithm is commonly used to address multi-class classification problems?
  • What is the primary focus of Bessel's correction in statistical calculations?
  • Which term is commonly associated with removing common words that may not add significant meaning in text analysis?
  • What is the name of the process that converts a continuous variable into a discrete variable?
  • What type of machine learning is characterized by using multiple layers of information to make complex decisions?
  • What technique improves a model's ability to generalize to new data by partitioning the data?
  • What defines a specific implementation of an algorithm that generates predictions based on training data?
  • What term describes massive quantities of data that cannot be easily translated into actionable intelligence using traditional methods?
  • What process can help in making raw data more understandable and usable?
  • Which of the following are examples of pipeline monitoring solutions?
  • What type of plot represents a probability distribution using bins?
  • Which methodology involves identifying, analyzing, and controlling variables in an experimental setup?
  • Which of the following is commonly used as a metric for constructing a decision tree?
  • Which term describes a distribution characterized by a flat peak and light tails?
  • What type of variable is often the focus of a regression analysis?
  • What algorithm is commonly used to classify data examples based on similarities within the feature space?
  • What method involves optimizing hyperparameters through random sampling of parameter combinations?
  • Which function outputs a value between 0 and 1, forming an S shape in logistic regression?
  • Which hyperparameter specifies how many samples are required to split a decision node?
  • Which analysis focuses on understanding and modeling the distribution and relationships of data points?
  • In k-fold cross-validation, how is the data used?
  • Which type of machine learning provides known label values as input for future predictions?
  • Which technique involves scaling features so that the mean value is 0 and the standard deviation is 1?
  • Which of the following is true about WCSS in clustering?
  • What type of plot visually represents data values using varying shades of color on a matrix?
  • What is a measure that indicates the strength of dependence between two variables, producing a value between +1 and -1?
  • What is the primary function of a bar chart?
  • What measures how often a learning model incorrectly classifies positive outcomes?
  • Which term is used to describe categorical values visualized for comparison purposes in a dataset?
  • What does data cleaning refer to?
  • What defines non-parametric algorithms?
  • What are the internal parameters derived from a model during the training process known as?
  • Which concept is used to quantify the error between estimated values and actual labeled values?
  • What clustering method adjusts the number of clusters dynamically by merging nearby points?
  • In clustering, what is the term for the point where the mean distance between data examples and their centroid stabilizes?
  • What does a z-score indicate?
  • Which term refers to the situation where the distribution of data has two peaks?
  • What does the ROC curve represent in model evaluation?
  • In data science, what is the purpose of data preprocessing?
  • What role does a dataset play in the business goals of a project?
  • What is the term for a type of data analysis that quantitatively summarizes patterns and relationships in a dataset?
  • Which metric measures the distance between a data point and its cluster centroid?
  • What transformation method raises each data example to a power of some lambda value to reduce skewness?
  • Which type of algorithms typically has a fixed number of parameters?
  • In leave-p-out validation, how many data points are used for testing?
  • What do we call irrelevant or irregular data values that obscure meaningful patterns in other relevant data?
  • Which method is used for evaluating model performance in regression tasks?
  • Which term describes a distribution with a specific shape characterized by a normal peak and normal tails?
  • Which term best describes the alterations made to data in order for it to support analytics?
  • In data science, what method is used when a data example can only be classified as a 1 or 0?
  • Which of the following statements is true about model parameters?
  • In sensitivity analysis, which aspect is primarily evaluated?
  • What is meant by collinearity in regression analysis?
  • In the context of text analysis, what is the primary purpose of a bag-of-words model?
  • Which of the following is not a characteristic of data visualizations?
  • What does IQR stand for in statistical analysis?
  • What type of data can be placed in an order?
  • Which AI discipline enables machines to improve estimative capabilities without explicit instructions?
  • What characteristic of a distribution indicates that values are concentrated toward one of the extremes?
  • What is a 'dataset' in the context of data science?
  • What term describes the frequency with which a machine learning model correctly identifies all actual negative instances?
  • What term describes the measure of decision-making processes in a model applied to specific data examples?
  • What issue occurs when a model is too complex and matches the training data too closely?
  • In machine learning, what is the variable that you are attempting to predict in a training set called?
  • Which characteristic is involved in a learning curve?
  • What term describes a distribution that follows the shape of a normal curve?
  • Which term describes a mathematical system that generates assumptions about data using statistical methods?
  • What process is crucial for making raw data interpretable and analyzable by machine learning algorithms?
  • Which term is used to indicate a data point's placement within a cluster in relation to others?
  • What technique helps prevent overfitting in machine learning by constraining model parameters?
  • Which type of distribution demonstrates the probability of outcomes for a random variable?
  • Data that can take any value within a range is known as which type of data?
  • What type of error cannot be reduced further when fitting a machine learning model due to model framing?
  • Which approach aims to identify and control the variables in an experiment?
  • What is the defining feature of parametric algorithms?
  • What best describes the primary goal of data science?
  • In the context of machine learning, the concept of "noise" can best be described as:
  • What does a 'decision boundary' separate in a dataset?
  • What is the concept of model drift?
  • What regulation governs the export of EU citizens' personal data?
  • What is a limitation of linear equations in data modeling?
  • What does 'data munging' or 'data wrangling' primarily involve?
  • How does a confidence interval typically perform in relation to the true population mean?
  • A value that falls outside the expected range of data can best be described as what?
  • What measure indicates the linear correlation between two variables commonly called x and y?
  • What does the term 'data wrangling' typically refer to in data science?
  • What is the primary goal of a regression analysis?
  • What key benefit does cross-validation provide in model evaluation?
  • What is the hyperparameter optimization method that evaluates multiple parameter combinations?
  • Which variable in an experiment is typically manipulated to observe its effect on the dependent variable?
  • What type of data is described as unstructured?
  • What does recall measure in a machine learning model?
  • What iterative ensemble learning method builds multiple decision trees to reduce errors?
  • What term describes a representation of the relationship between input and output variables in a machine learning model?
  • What type of data is defined as being in a format that facilitates searching, filtering, or extracting?
  • What is the formula for calculating accuracy in a classification model?
  • What is the name of the process used to fill in missing data values using statistical calculations?
  • Which clustering method is noted for its shortcomings with circular or spiral data?
  • What is the key benefit of using DevOps practices in data science projects?
  • What type of distribution is characterized by having two humps?
  • In a confusion matrix, what do true positives represent?
  • Which process converts data from one type to a coded value of a different type?
  • Which of the following is a common method for handling missing data?
  • What is the method for systematically designing experiments to evaluate the influence of variables?
  • What is the classification algorithm that utilizes Bayes' theorem to compute classification probabilities called?
  • Which statistical measures summarize the "middle" portion of a sample dataset?
  • Which clustering algorithm starts with all data examples in a single cluster and splits them?
  • What does leave-one-out validation involve?
  • What is the term used to describe a project's uncontrolled growth beyond its original objectives?
  • The CART model is primarily used for which of the following?
  • Which method minimizes a cost function by gradually tuning model parameters?
  • In machine learning, what is the significance of the term 'errors'?
  • What formula is used to calculate the variance of a population?
  • What is the term for the process of simplifying a decision tree by removing nodes, branches, and leaves that offer little value?
  • What type of data can be easily searched and filtered, while other elements are not?
  • Which ensemble learning method aggregates multiple decision tree models together and selects the optimal classifier or predictor?
  • What does the p-value represent in hypothesis testing?
  • What is lemmatization primarily used for in text processing?
  • In the context of data science, what is the purpose of a cost function?
  • Which term refers to qualitative data with a limited number of values?
  • Which of the following is a method used in hypothesis testing?
  • What is defined as a value that deviates significantly from the main distribution of values?
  • Which statistical parameter is NOT part of the common four used to measure distributions?
  • What is the process of identifying an issue that should be addressed and putting it in understandable and actionable terms called?
  • What term describes a type of classification in SVMs where all examples are on the correct side of the margin?
  • What does the abbreviation AI stand for in data science?
  • Which of the following is a technique for selecting a subset of features in a model?
  • Which regression method is often used for modeling binary outcomes?
  • Which term encompasses methods like linear regression, decision trees, and k-means clustering?
  • Which of the following structures is characterized by conditional statements and their conclusions?
  • Which of the following techniques is used to address class imbalance in datasets?
  • What does R² indicate in statistical modeling?
  • What does TPR stand for in the evaluation of machine learning models?
  • The absence of reported observations in the training data may lead to which type of bias?
  • Which process is essential for ensuring data quality before analysis?
  • What property indicates that a process cannot perfectly estimate individual events but can demonstrate a general pattern?
  • Which clustering algorithm starts with each data example in its own cluster?
  • What does LCA stand for in data analysis?
  • Which approach combines the estimates of multiple models in machine learning?
  • In the context of data science, what does 'data preparation' specifically refer to?
  • Which chart type is used to represent the proportional measurement of categorical values with horizontal or vertical bars?
  • What aspect of data is preserved in stratified k-fold cross-validation?
  • Which of the following sampling techniques could lead to an underrepresentation of minority classes?
  • Which regression technique forces the coefficients of the last relevant feature to zero using the l1 norm?
  • What is the challenge of 'data wrangling' primarily about?
  • What can be inferred if kurtosis is less than 3?
  • What is the expected output when applying standardization to a dataset?
  • In machine learning, what type of classification problem allows data examples to be classified into one of three or more classes?
Subscribe

Get the latest from Examzify

You can unsubscribe at any time. Read our privacy policy