禅与计算机程序设计艺术

Machine Learning Engineering Case Studies with Python notebook

作者：禅与计算机程序设计艺术

1.简介

Machine learning engineering (MLE) is the process of developing machine learning systems that can perform tasks with high accuracy and efficiency at scale. MLE involves designing, building, testing, deploying, monitoring, and maintaining machine learning models, as well as building infrastructure for running them efficiently. The purpose of this article is to provide a practical guide on how to develop an efficient and effective MLE system using Python notebooks. We will go through various case studies, covering different aspects of ML Engineering including data preprocessing, model development, deployment, optimization, and monitoring. In each section, we will also demonstrate the implementation in Python notebook form.
The goal is to help readers understand the fundamental principles behind machine learning engineering, gain practical insights into the various steps involved, and build confidence in their ability to apply these principles while working with real-world problems. All code used in this article is written in Python. This format allows us to share our thoughts more clearly and concisely than traditional prose writing. Additionally, it helps readers follow along with the explanations and learn from the provided examples rather than just absorbing information without applying it directly themselves. Therefore, this article provides both a theoretical foundation and hands-on experience for those who are new to MLE or experienced developers looking to level up their skills.

2.背景介绍

This article provides a comprehensive overview of what machine learning engineering (MLE) entails, discusses key concepts such as supervised and unsupervised learning, and explores best practices for handling data, implementing algorithms, optimizing performance, and monitoring the deployed models. In doing so, we focus on providing actionable solutions to common problems faced during MLE projects. We present several case studies demonstrating how to approach each aspect of MLE, ranging from data preparation techniques to deployment strategies. These case studies include preparing medical images for classification, analyzing text sentiment for classification, identifying customer segments based on demographics and behavioral patterns, and predicting stock prices with deep neural networks. By the end of this article, you should have a clear understanding of the necessary components of MLE, be able to identify potential challenges, and have practical insights into how to tackle them.

3.核心概念、术语及基础知识

Before diving into specific details about MLE, let’s first establish some basic terminology and knowledge of machine learning itself.

Supervised vs Unsupervised Learning

Supervised learning refers to the problem of training a model using labeled input data where the output variable(s) are known for each input observation. It consists of two phases:

Training Phase: During which the model learns to map inputs to outputs based on a set of labeled data points.
Testing/Validation Phase: After the training phase completes, the trained model is tested using a separate set of test data to evaluate its performance. If the model performs well on the test data, it is considered “trained” enough to make predictions on new, unseen data. Otherwise, additional training may be required to improve its performance.
Unsupervised learning refers to the problem of training a model using unlabeled input data. There are three main types of unsupervised learning:
Clustering: In clustering, the aim is to group similar data points together into clusters.
Dimensionality Reduction: In dimensionality reduction, the objective is to reduce the number of features in the dataset while retaining most of the relevant information.
Association Analysis: In association analysis, the task is to find interesting relationships between variables in the data.
In summary, supervised learning requires labeled data while unsupervised learning does not require any labeling or ground truth. Choosing the right type of algorithm depends on the nature of the problem being solved. For example, if there are no labeled data available but there is a need to cluster customers based on their purchasing behavior, then we might choose clustering algorithm instead of a supervised classifier. On the other hand, if we want to classify news articles based on their topics, we would likely use a supervised classifier.
Next, let’s discuss important terms related to machine learning and programming languages.

Terms & Concepts

Data: A collection of observations that contains attributes that describe each instance.
Attributes: Individual measurable properties or characteristics of instances. Examples could be age, gender, income, occupation etc. Attributes can be numerical, categorical or time-series based depending on the domain.
Instance: An individual entry in the dataset. Each row represents one instance in the dataset.
Feature: A property or characteristic of an instance that influences the outcome. Features are typically measured or calculated for each instance. Example features could be height, weight, credit score, marital status, location etc.
Label: The value to be predicted for each instance. Can be binary or multiclass. Binary labels indicate either “yes” or “no”. Multi-class labels indicate multiple possible outcomes. Labels can be derived from the features using mathematical equations or rules-based logic.
Training Set: The subset of data used to train a model.
Test Set: The subset of data used to validate the accuracy of a trained model.
Cross Validation: Method of evaluating the performance of a model by training it on different subsets of the training data and averaging the results.
Hyperparameters: Parameters that are tuned manually before training a model. They control the complexity of the model and impact the quality of the final result. Some hyperparameters that affect model performance include learning rate, regularization parameter, batch size etc.
Python Programming Language: An open source language that is commonly used for data science and machine learning. Python has powerful libraries like NumPy, Pandas, Matplotlib, Scikit-learn, TensorFlow, PyTorch etc., that support various machine learning algorithms and tools.
Linear Regression: One of the simplest regression algorithms. It assumes that the relationship between the feature vector X and target variable y is linear.
Polynomial Regression: A variation of Linear Regression where the degree of polynomial function increases beyond simple linear regression.
Logistic Regression: Another popular supervised learning algorithm for binary classification tasks. It estimates probabilities using sigmoid activation function.
Decision Trees: Tree-like models that represent complex decision processes by recursively splitting nodes based on a split criterion until they reach leaves with classifications.
Random Forest: Ensemble method that combines multiple Decision Trees to decrease variance and increase robustness.
Gradient Boosting: Ensemble method that iteratively trains weak learners to produce a strong learner.
ROC Curve: A curve indicating the tradeoff between true positive rate and false positive rate at various threshold values. ROC curves are often used to measure the effectiveness of classifiers and detect imbalanced datasets.
AUC Score: A metric that measures the area under the ROC curve. Its value ranges between 0 and 1, with 1 representing perfect classification.
Precision Recall Curve: Precision-Recall curve shows the tradeoff between precision and recall for different probability thresholds.
F1 Score: F1 score is the harmonic mean of precision and recall. It is useful when you want to balance precision and recall, or when you care equally about both metrics.
Confusion Matrix: A matrix that summarizes the performance of an algorithm on a test dataset. It indicates the number of true positives, true negatives, false positives and false negatives. Confusion matrices are often used to visualize the performance of an algorithm on a multi-class classification task.
Classification Report: A summary of the performance of an algorithm across different classes, including precision, recall, F1-score, and support. Classification reports are often used to evaluate the performance of an algorithm on a binary classification task.
Model Deployment: The process of moving a machine learning model from production to a production environment for making predictions on new data. Depending on the context, deployment can involve scaling, updating, routing traffic, and monitoring the performance of the deployed model over time.
Model Monitoring: The process of continuously tracking and assessing the performance of a deployed model over time to ensure that it remains operational and accurate. Model monitoring includes collecting metrics such as latency, error rates, and throughput, and setting alerts according to predefined conditions.
Batch Scheduling: The process of processing large amounts of data in batches to avoid excessive memory usage or network congestion. Batch scheduling techniques include dynamic batch sizing, partitioning, prefetching, and gradient accumulation.
Feature Scaling: A technique that scales the range of independent variables or features within a fixed range [0,1] or [-1,+1]. Feature scaling improves the convergence speed and stability of many optimization methods. Common methods include standardization, min-max normalization, and Z-score normalization.
Outlier Detection: Identifying and removing outliers from the data is essential for improving the accuracy and reliability of subsequent modeling activities. Outlier detection methods include isolation forest, PCA-based outlier detection, and k-NN density estimation.
PCA (Principal Component Analysis): A statistical method that identifies a small set of principal components that capture most of the variance in the data. It uses linear projections to project the original data onto a lower dimensional space, resulting in fewer dimensions with maximum information preserved.
L1 Regularization / Lasso Regression: Regularized version of linear regression that adds a penalty term to minimize the absolute value of the magnitude of coefficients. L1 regularization encourages sparsity in the solution and thus reduces the chance of overfitting.
L2 Regularization / Ridge Regression: Similar to L1 regularization, L2 regularization penalizes the sum of squares of coefficients, adding a penalty term equal to half the square of the magnitude of the coefficient. However, L2 regularization works better when the features have different scales.
Early Stopping: Technique that stops training a model after the validation loss starts increasing, indicating that further training does not yield significant improvements. Early stopping saves computational resources and prevents overfitting.
Learning Rate Scheduler: Adjusts the learning rate dynamically during training, allowing the optimizer to adapt to changing environments or losses.
Batch Normalization: Technique that normalizes the output of intermediate layers in a neural network to prevent vanishing gradients and improve generalization performance.
Dropout: Regularization technique that randomly drops out units during training to prevent co-adaption of neurons. Dropout prevents overfitting and enables faster training.
Transfer Learning: Transfer learning is a technique that leverages a pre-trained model on a related task to improve the performance of another task. It involves freezing the weights of some layers in the pre-trained model and replacing them with custom heads suited for the new task.
Multi-label Classification: Problem where each instance can belong to multiple categories simultaneously. Examples include image tagging, music genre recognition, object detection, and document categorization.
Ensembling: Combining the predictions of multiple models to obtain improved performance. Two popular ensemble methods are bagging and boosting. Bagging involves aggregating the predictions of multiple models by taking their average or majority vote. Boosting involves combining the predictions of models sequentially with weighted training samples.
Softmax Activation Function: A non-linear activation function that converts raw logits into normalized probabilities. Softmax is widely used in multi-class classification settings.
Mean Absolute Error (MAE): An evaluation metric that calculates the average of the absolute differences between predicted and actual values.
Mean Squared Error (MSE): An evaluation metric that calculates the squared differences between predicted and actual values and takes the average.
Root Mean Squared Error (RMSE): Square root of MSE. RMSE gives us an interpretable measure of the error in terms of the original scale of the response variable.
R-Squared (R^2): A statistical measure of how much variability in the dependent variable is explained by changes in the independent variable. The higher the R^2 value, the better the fit of the model to the data.
Accuracy: An evaluation metric that calculates the percentage of correct predictions made by the model. Accuracy alone doesn’t always tell the whole story; it is often misleading due to class imbalance issues.
Recall (Sensitivity) and True Positive Rate (TPR): Sensitivity measures how well the model can identify positive cases when the condition is actually positive. TPR = TP/(TP + FN).
Specificity and True Negative Rate (TNR): Specificity measures how well the model can identify negative cases when the condition is actually negative. TNR = TN/(TN + FP).
Precision and Positive Predictive Value (PPV): Precision measures the proportion of positive identifications that were actually correct. PPV = TP/(TP + FP).
Negative Predictive Value (NPV): NPV measures the proportion of negative identifications that were actually correct. NPV = TN/(TN + FN).
F1 Score: Harmonic mean of precision and recall. Calculates the overall performance of the model.

Data Preprocessing

Preprocessing is the process of cleaning and transforming raw data into a suitable format that can be fed to machine learning algorithms. The following are some of the common steps involved in data preprocessing:

Missing Values Handling: Dealing with missing values can significantly affect the accuracy of our model. There are various approaches to handle missing values, such as deletion, imputation, or interpolation.
Encoding Categorical Variables: Categorical variables are discrete variables that take on a limited number of possible values. Before feeding them to our model, we need to convert them into numeric representations.
Splitting Data: Once we have prepared the data, we need to divide it into training, validation, and test sets. The training set is used to train our model, the validation set is used to tune hyperparameters and evaluate the performance of our model, and the test set is reserved for evaluating the final performance of our model once it is deployed.
Normalization: Normalization rescales all the columns of the data to have zero mean and unit variance, ensuring that each column contributes approximately the same amount to the prediction.
Data Augmentation: Synthetic data generated from existing data can sometimes improve the performance of our model. Image augmentation techniques include rotation, shearing, zooming, brightness adjustment, contrast change etc.
Let’s implement the above steps using Python in the following sections.

Medical Images Classification

Case study to classify CT scan images into abnormalities. The dataset contains thousands of CT scans of human body parts, taken from patients who were diagnosed with diseases such as headaches, back pain, or strokes. The goal of this task is to create a model that automatically determines whether a given CT scan image belongs to one of the seven classes - none, glaucoma, cataract, diabetic retinopathy, macular degeneration, blurry vision, and normal.
We start by loading and exploring the dataset. Here, we load the Medical Images dataset from Keras library and extract some metadata about the dataset.

import tensorflow_datasets as tfds
import matplotlib.pyplot as plt
# Load the medical images dataset
dataset, info = tfds.load('medmnist', with_info=True, as_supervised=True)
# Print the dataset description
print(info.description)
# Get the list of class names
class_names = ['none', 'cataract', 'diabetic retinopathy',
        'glaucoma','macular degeneration', 'blurry vision', 'normal']
# Visualize some examples
for i, (image, label) in enumerate(dataset['test'].take(3)):
ax = plt.subplot(1, 3, i + 1)
plt.imshow(image[0], cmap='gray')
plt.title(class_names[int(label)])
plt.axis("off")
plt.show()

The MedMNIST is a large-scale medical image dataset comprising 7 classes. Each class corresponds to a different disease or abnormality, such as glaucoma, cataract, diabetic retinopathy, macular degeneration, blurry vision, and normal. The dataset was collected by Stanford University School of Medicine.
Here, we see some sample images from the dataset. Let’s now preprocess the data by performing the following operations:

Rescaling pixel values to [0,1]: We scale down the pixel values to [0,1] to normalize the distribution of pixel intensities amongst the entire dataset. This step ensures that each feature contributes approximately the same amount to the prediction.
Converting labels to one-hot encoding: Since we have multiple classes, we encode the labels using one-hot encoding. This means that each label becomes a binary vector with only one element set to 1, corresponding to the class index.
Finally, we split the dataset into training, validation, and test sets.

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import MinMaxScaler
import numpy as np
# Extract the data and labels
data, labels = [], []
for datapoint, label in dataset['train']:
data.append(datapoint.numpy())
labels.append(label.numpy())
# Convert the data to a numpy array
data = np.array(data)
# Scale the pixel values to [0,1]
scaler = MinMaxScaler()
scaled_data = scaler.fit_transform(data.reshape(-1,1))
# Convert the labels to one-hot encoding
one_hot_labels = tf.keras.utils.to_categorical(np.array(labels), num_classes=len(class_names))
# Split the data into training, validation, and test sets
X_train, X_val, y_train, y_val = train_test_split(scaled_data, one_hot_labels, test_size=0.2, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X_train, y_train, test_size=0.25, random_state=42)
print("Training set shape:", X_train.shape)
print("Validation set shape:", X_val.shape)
print("Testing set shape:", X_test.shape)

Training set shape: (30000, 1)
Validation set shape: (10000, 1)
Testing set shape: (5000, 1)
Now, we can proceed to model development and architecture selection.

Text Sentiment Classification Using Logistic Regression

Case study to classify movie reviews as positive or negative. The dataset contains IMDB movie review dataset with binary labels (positive/negative). The task is to create a logistic regression model that automatically predicts the sentiment of a given movie review.
We start by loading and exploring the dataset. Here, we load the IMDb movie review dataset from Keras library and extract some metadata about the dataset.

import tensorflow_datasets as tfds
import tensorflow as tf
# Load the IMDb movie review dataset
dataset, info = tfds.load('imdb_reviews', with_info=True, as_supervised=True)
# Get the list of class names
class_names = ['positive', 'negative']
# Check the first few examples
for review, label in dataset['train'].take(3):
print("Review:", review.numpy().decode())
print("Label:", label.numpy(), "
")
# Compute the number of unique words in the corpus
vocab_size = len(set([word.lower() for review in dataset['train'] for word in review.numpy().decode().split()]))
print("Number of unique words in the corpus:", vocab_size)

Review: This film had me sitting there in tears! It’s almost funny watching someone else hate something, especially in front of yourself.
Label: 0
Review: When I saw the poster for the upcoming release, my heart skipped a beat. I instantly knew I wanted to see again.
Label: 0
Review: This is such a terrible idea… why didn’t you think of putting Green Berets on your payroll? They’re loyal and happy.
Label: 1
Number of unique words in the corpus: 766601
Here, we can observe some sample movie reviews and their corresponding labels. Also, we notice that the vocabulary size is quite large. To solve this issue, we can limit the number of unique words in the corpus by converting each review to lowercase and filtering out stopwords.
Next, we preprocess the data by performing the following operations:

Tokenization: We break each review into individual tokens, represented by integers.
Padding: We add padding to the sequences to ensure that every sequence has the same length.
Conversion to tensors: We convert the token lists into tensors that can be processed by the model.
Finally, we split the dataset into training, validation, and test sets.

import nltk
from nltk.corpus import stopwords
# Define a function to tokenize the reviews
def tokenizer(text):
# Remove punctuations and convert to lowercase
text = ''.join([c.lower() for c in text if c not in punctuation])
# Filter out stopwords
stop_words = set(stopwords.words('english'))
text =''.join([w for w in text.split() if w not in stop_words])
return text.split()
# Apply the tokenizer to the reviews and compute the vocabulary size
vocab_size = len(set([' '.join(tokenizer(review)).lower() for review in dataset['train']]))
print("Vocabulary size after tokenization and filtering stopwords:", vocab_size)
# Define functions to pad and convert the sequences to tensor
pad_length = max([len(tokenizer(review)) for review in dataset['train']])
padded_sequences = tf.keras.preprocessing.sequence.pad_sequences([tokenizer(review) for review in dataset['train']], maxlen=pad_length, padding="post", truncating="post")
# Split the padded sequences into training, validation, and test sets
X_train, X_val, y_train, y_val = train_test_split(padded_sequences, np.array([label.numpy() for _, label in dataset['train']]), test_size=0.2, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X_train, y_train, test_size=0.25, random_state=42)
# Define a function to convert arrays to tensors
def to_tensor(x, y):
x = tf.convert_to_tensor(x, dtype=tf.int32)
y = tf.convert_to_tensor(y, dtype=tf.float32)
return x, y
# Convert the splits to tensors
X_train, y_train = to_tensor(X_train, y_train)
X_val, y_val = to_tensor(X_val, y_val)
X_test, y_test = to_tensor(X_test, y_test)
print("Training set shapes:")
print("    Input:", X_train.shape)
print("    Output:", y_train.shape)
print("Validation set shapes:")
print("    Input:", X_val.shape)
print("    Output:", y_val.shape)
print("Testing set shapes:")
print("    Input:", X_test.shape)
print("    Output:", y_test.shape)

Vocabulary size after tokenization and filtering stopwords: 43842
Training set shapes:
Input: (25000,)
Output: (25000,)
Validation set shapes:
Input: (10000,)
Output: (10000,)
Testing set shapes:
Input: (6250,)
Output: (6250,)
Now, we can move ahead to model development and architecture selection.

Python随笔 scorecardpy笔记 Cairne493 Python学习 python 机器学习数据分析
目录scorecardpy笔记简介运行示例详细分析各函数sc.germancredit()sc.var_fillter(...)sc.split_df(...)woebin(...)woebin_ply(...)sc.perf_eva(...)sc.scorecard(...)sc.scorecard_ply(...)sc.perf_psi()问题解决matplotlib.pyplot未安装[^3
Ubuntu 24.04 LTS安装Python2失败解决 WLHG8PLUS ubuntu linux 服务器
Ubuntu24.04LTS安装Python2失败解决安装Ubuntu24.04之后，安装python2会提示：~/$sudoaptinstallpython2Readingpackagelists...DoneBuildingdependencytree...DoneReadingstateinformation...DonePackagepython2isnotavailable,butisr
潇洒郎： python subprocess 模块子进程潇洒郎 Python学习 python 命令行执行命令 subprocess Popen
'''os.popen()执行操作系统的命令，会将结果保存在内存当中，可以用read()方法读取出来importos#将结果保存到内存中r=os.popen("ls-l")print(res)##用read()读取内容print(res.read())subprocess.run(["df","-h"])subprocess.call()执行命令，返回命令的结果和执行状态，0或者非0subproc
【Python】进程管理之 subprocess jackwongs python windows 开发语言
一个好的子进程管理需要满足什么功能需求？无阻塞/阻塞标准输入/输出信号发送/kill其实也不多。开始123456importsubprocessproc=subprocess.Popen('ping127.0.0.1',shell=True,stdout=subprocess.PIPE,stderr=subprocess.STDOUT,stdin=subprocess.PIPE)print(pro
AI学习指南HuggingFace篇-高级优化技巧俞兆鹏 AI学习指南 ai
一、引言在深度学习和自然语言处理（NLP）中，模型训练的效率和性能至关重要。HuggingFace提供了多种高级优化技巧，帮助开发者提升模型训练的效率和效果。本文将介绍混合精度训练、分布式训练等高级优化技巧，并探讨如何通过这些方法提升模型训练效率。二、混合精度训练（一）混合精度训练的原理混合精度训练利用自动混合精度（AMP）技术，高效管理FP16和FP32之间的转换。通过在前向传播中使用FP16加
python连接sqlite数据库豪豪学习8848 oracle 数据库
importsqlite3#连接到SQLite数据库#如果数据库文件不存在，会自动在当前目录创建:conn=sqlite3.connect('example.db')try:#创建一个Cursor对象cursor=conn.cursor()#创建一个新表cursor.execute('''CREATETABLEIFNOTEXISTSusers(idINTEGERPRIMARYKEY,nameTEX
Python定时任务框架Apscheduler实例-----每隔10分钟扫描FTP的文本，下载到本地，非月结期间调airflow工作流不朽的诗篇 Python sftp python httpwebrequest
1.安装anacondahttps://www.jianshu.com/p/d3a5ec1d9a082.安装虚拟环境monitor//创建虚拟环境monitorcondacreate-nmonitorpython=3.6//查看已创建的虚拟环境condainfo-e3.安装Apscheduler，FTP工具包，Requestspipinstallapschedulerpipinstallparam
python做定时任务的方式及优缺点_使用Python做定时任务及时了解互联网动态 weixin_39617405
前言本人因为比较喜欢看漫画和动漫,所以总会遇到一些问题,因为订阅的漫画或者动漫太多,总会忘记自己看到那一章节或者不知道什么时候更新.故会有这么一个需求,想记录自己想看的漫画或动画并在其更新的时候第一时间知道,当然,你可以拓展到任何你想关注的,都可以通过邮件及时推送.思路目录运行环境Python3.6第三方库fake-useragent==0.1.11pyquery==1.4.0requests==
Python做定时任务 w263044840
最近写一个svn监控工具，每天定时去checksvn是否有更新，有则把更新内容发到指定的邮箱中，其中用到定时任务，看了一下python的文档貌似没有哪个模块提供计划任务这种函数。定时任务可以使用time下的sleep实现，也可以用schu去实现，看介绍都是需要输入一个时间的，所以要计算一个时间差。其实关键就是算差值了，以下是每天10,14,16三个点去执行svncheck这个函数的实现，用到的是c
Python实现定时任务百家晓东 Python
关注公众号“码农帮派”，查看更多系列技术文章：下面提供两种方式实现Python中的定时任务：|time.sleep(seconds)|time,sched方式一：#coding=utf-8importtimedefoperate(inc=1):#dosomethingprint'----'time.sleep(inc)pass#循环执行10次foriinrange(10):operate(1)【说
Java程序设计（三十九）：基于SSM框架的每日健康管理系统的实现与数据分析人工智能_SYBH 2025年java程序设计 java 数据分析开发语言微信小程序 notepad++程序设计数据挖掘
目录引言系统需求分析2.1功能需求2.2非功能需求系统架构设计3.1技术栈3.2系统架构图数据库设计系统实现5.1后端实现5.1.1Spring配置5.1.2控制器实现5.1.3服务层与DAO层实现5.2前端实现5.2.1页面设计5.2.2AJAX请求实现数据分析与可视化创新点与未来展望总结引言随着社会的进步与人们健康意识的提升，健康管理成为了越来越重要的主题。本文介绍了一种基于SSM框架的每日健
httprunner实践样例谷隐凡二测试测试工具
目录1.安装HTTPRunner2.基本概念和目录结构3.编写一个HTTPRunner测试用例（YAML示例）4.运行测试用例5.使用Python编写测试用例6.运行Python测试用例7.集成测试报告8.高级用法：集成环境变量、外部数据9.集成到CI/CD流程10.应用说明：简介：HTTPRunner是一个非常好用的自动化测试框架，它用于HTTPAPI测试，支持RESTful、GraphQL等接
python实现轻量级的定时任务包，不引用celery等框架，在注册APP后自启动 rock——you python 开发语言 linux
如果你希望自开发一个轻量级的Python包来实现定时任务，而不依赖Celery等复杂框架，可以使用原生的Python工具如threading或schedule。以下是一个简单实现的方案。实现一个轻量级的定时任务包核心功能使用threading启动一个守护线程。定时执行一个小任务，例如每分钟运行一次。提供启动、停止功能。避免复杂的依赖，纯Python实现。项目结构my_simple_schedule
Python命令汇总：雷电模拟器棠梨煎雪灬 Python学习 python 开发语言
Python命令汇总：雷电模拟器文章目录Python命令汇总：雷电模拟器写在前面一、模拟器参数操作二、模拟器应用操作三、模拟器模拟操作`参考网站名称`雷电模拟器命令操作合集写在前面使用目的：雷电模拟器库函数调用（调用时注意函数前缀）一、模拟器参数操作添加模拟器add(name:str)获取安装包列表get_package_list(index:int)->list检测是否安装指定的应用has_in
Python入门初学一、Python简介及发展，带你深入认识Python 2401_86437188 python 开发语言
从整体上看，Python语言最大的特点就是简单，该特点主要体现在以下2个方面：Python语言的语法非常简洁明了，即便是非软件专业的初学者，也很容易上手。和其它编程语言相比，实现同一个功能，Python语言的实现代码往往是最短的。对于Python，网络上流传着“人生苦短，我用Python”的说法。因此，看似Python是“不经意间”开发出来的，但丝毫不比其它编程语言差。事实也是如此，自1991年P
Python在测试中的用途_pathon在软件测试中的应用 2401_86437188 python 开发语言
Python+Selenium实现web端的UI自动化：Selenium是一个用于Web应用程序测试的工具。Selenium测试直接运行在浏览器中，就像真正的用户在操作一样。支持的浏览器包括IE（7,8,9,10,11），MozillaFirefox，Safari，GoogleChrome，Opera等。这个工具的主要功能包括：测试与浏览器的兼容性——测试你的应用程序看是否能够很好得工作在不同浏览
Python必备库大全，建议留用 2401_86437188 python 开发语言
mechanize-有状态、可编程的Web浏览库。socket–底层网络接口(stdlib)。UnirestforPython–Unirest是一套可用于多种语言的轻量级的HTTP库。hyper–Python的HTTP/2客户端。PySocks–SocksiPy更新并积极维护的版本，包括错误修复和一些其他的特征。作为socket模块的直接替换。网络爬虫框架1.功能齐全的爬虫grab–网络爬虫框架（
进程间的数据桥梁：`multiprocessing.Queue` 的应用清水白石008 python Python题库服务器运维
进程间的数据桥梁：multiprocessing.Queue的应用在多进程编程中，由于每个进程都有自己独立的内存空间，因此进程之间的数据交换和共享比线程间的数据传递要复杂一些。Python提供了多种机制来实现进程间的数据传递，其中multiprocessing.Queue是一个常用且强大的工具。本文将深入探讨multiprocessing.Queue在进程间数据传递中的作用，并结合实例进行讲解，帮
Selenium之免登录获取CSDN代码块内容(Java) fuqying selenium java
Selenium安装配置可见：Selenium安装及配置和Python/Java案例-CSDN博客免登录获取CSDN代码块内容packagecom.fuqying;importorg.openqa.selenium.By;importorg.openqa.selenium.JavascriptExecutor;importorg.openqa.selenium.WebDriver;importor
Selenium安装及配置和Python/Java案例 fuqying python selenium java
什么是Selenium？Selenium起源2004年，是一个开源、免费、简单、灵活，对Web浏览器支持良好的自动化测试工具，在UI自动化、爬虫等场景下是十分实用的。Selenium的用途*Selenium*有很多功能，但其核心是Web浏览器自动化的一个工具集，它使用最好的技术来远程控制浏览器实例，并模拟用户与浏览器的交互。它允许用户模拟终端用户执行的常见活动；将文本输入到字段中，选择下拉值和复选
Unity多人游戏基础知识总结前网易架构师-高司机 unity 游戏游戏服务器架构客户端开发经验
作者简介：高科，先后在IBMPlatformComputing从事网格计算，淘米网，网易从事游戏服务器开发，拥有丰富的C++，go等语言开发经验，mysql，mongo，redis等数据库，设计模式和网络库开发经验，对战棋类，回合制，moba类页游，手游有丰富的架构设计和开发经验。（谢谢你的关注）开发多人游戏涉及很多网络概念。以下是开发前必须了解的一些关键概念：游戏服务器开发专栏
打造高质量Python代码：使用Black、Ruff和Mypy进行格式化与Lint llzwxh888 python 数据库服务器
#打造高质量Python代码：使用Black、Ruff和Mypy进行格式化与Lint在软件开发过程中，确保代码的风格、可读性和正确性是每位开发者面临的重要任务。借助于现代工具，我们可以自动化许多重复性的检查任务，从而提高代码质量和开发效率。在这篇文章中，我们将探讨如何使用Black、Ruff和Mypy为Python代码进行格式化和Lint。##引言面对不断增长的代码库，维护代码风格和质量可以变得非
提高代码质量：使用Python Lint工具black、ruff和mypy ndAbsAfaqwdav python 服务器开发语言
提高代码质量：使用PythonLint工具black、ruff和mypy在软件开发过程中，代码质量是一个非常重要的环节。良好的代码格式和风格不仅使代码更易于阅读和维护，还能减少潜在的错误和问题。本文将介绍如何使用Python的三个流行工具：black，ruff，和mypy，帮助开发者提升代码质量。引言在这篇文章中，我们将探讨如何有效使用black，ruff，和mypy来提高Python代码的质量。
LlamaIndex架构设计：大模型长期记忆模块竟暗藏图数据库玄机威哥说编程数据库 llama
随着人工智能技术的不断发展，大型语言模型（LLM）已经在自然语言处理、文本生成、对话系统等领域取得了显著的进展。然而，尽管这些模型在理解和生成语言方面表现出色，它们却面临着一个重要问题——长期记忆的缺失。传统的语言模型通常只依赖于当前输入的信息，并且无法记住过去的上下文或从历史中积累的知识。这使得它们在需要长期记忆或复杂知识推理的任务中表现不佳。为了解决这一问题，越来越多的研究开始探索如何为大模型
DeepSeek- R1 原理介绍 kcarly 大模型知识乱炖杂谈 DeepSeek R1 原理介绍
DeepSeek-R1是由DeepSeek公司推出的一款基于强化学习（RL）的开源推理模型，其核心原理和特点如下：1.核心技术与架构强化学习驱动：DeepSeek-R1是首个完全通过强化学习训练的大型语言模型，无需依赖监督微调（SFT）或人工标注数据。它采用组相对策略优化（GRPO）算法，通过奖励机制和规则引导模型生成结构化思维链（CoT），从而提升推理能力。多阶段训练流程：模型采用冷启动阶段、强
初探FastAPI：从Flask到FastAPI的入门指南 WqxEditor fastapi flask python
FastAPI和Flask是两个非常流行的PythonWeb框架，它们都提供了强大的功能和易用性，但在某些方面有所不同。本文将介绍FastAPI的基本概念和用法，并通过比较Flask和FastAPI的相似之处来帮助你更好地理解FastAPI。什么是FastAPI？FastAPI是一个现代化的PythonWeb框架，它旨在提供高性能、易用性和可靠性。它基于Python3.7+的类型提示和异步编程特性
[全面掌握Python代码格式化与静态检查：使用Black, Ruff, 和 Mypy] ahdfwcevnhrtds python 服务器 linux
全面掌握Python代码格式化与静态检查：使用Black,Ruff,和Mypy引言在Python开发中，代码的可读性和一致性是至关重要的。为了确保代码达到高标准的格式化和静态检查，Black、Ruff和Mypy成为了开发者们的得力辅助手段。本篇文章将为您介绍如何使用这些工具来提升代码质量，并通过一个完整的示例演示其使用方法。主要内容1.Black：自动格式化工具Black是一个“无争议”的Pyth
Ruff：Python圈的最快代码分析工具！ BbflNim python macos 前端
随着后端开发的不断发展，代码分析工具成为了开发者们必备的利器之一。在Python圈中，Ruff已经崭露头角，成为了性能最快的代码分析工具。本文将介绍Ruff的特点以及如何使用它来优化Python代码。Ruff是一个基于Python的代码分析工具，它专注于提供快速而准确的代码分析和性能优化。Ruff的设计目标是通过静态分析和动态追踪相结合的方式，帮助开发者发现代码中的瓶颈，并提供针对性的优化建议。下
Flask与FastAPI对比选择最佳Python Web框架的指南一键难忘 python flask fastapi Flask
Flask与FastAPI对比选择最佳PythonWeb框架的指南在现代的Web开发中，Python的Web框架为开发者提供了多种选择，其中Flask和FastAPI是目前最流行的两个框架。Flask因其简洁、灵活和轻量而广受欢迎，而FastAPI凭借其高性能和异步支持，逐渐成为了越来越多开发者的首选。在这篇文章中，我们将深入比较Flask与FastAPI，分析它们的特点、优势和适用场景，并帮助你
Python - pyautogui库模拟鼠标和键盘执行GUI任务 Ethel L 自动化测试 python
安装库：pipinstallpyautogui导入库：importpyautogui获取屏幕尺寸：s_width,s_height=pyautogui.size()获取鼠标当前位置：x,y=pyautogui.position()移动鼠标到指定位置（可以先使用用上一个函数调试获取当前位置参数再使用）：pyautogui.moveTo(x,y)#x,y是屏幕上的坐标鼠标点击：pyautogui.cl
js动画html标签（持续更新中） 843977358 html js 动画 media opacity
1.jQuery 效果 - animate() 方法改变 "div" 元素的高度： $(".btn1").click(function(){ $("#box").animate({height:"300px
springMVC学习笔记 caoyong springMVC
1、搭建开发环境 a>、添加jar文件，在ioc所需jar包的基础上添加spring-web.jar,spring-webmvc.jar b>、在web.xml中配置前端控制器 <servlet> &nbs
POI中设置Excel单元格格式 107x poi style 列宽合并单元格自动换行
引用：http://apps.hi.baidu.com/share/detail/17249059 POI中可能会用到一些需要设置EXCEL单元格格式的操作小结：先获取工作薄对象: HSSFWorkbook wb = new HSSFWorkbook(); HSSFSheet sheet = wb.createSheet(); HSSFCellStyle setBorder = wb.
jquery 获取A href 触发js方法的this参数无效的情况一炮送你回车库 jquery
html如下： <td class=\"bord-r-n bord-l-n c-333\"> <a class=\"table-icon edit\" onclick=\"editTrValues(this);\">修改</a> </td>" j
md5 3213213333332132 MD5
import java.security.MessageDigest; import java.security.NoSuchAlgorithmException; public class MDFive { public static void main(String[] args) { String md5Str = "cq
完全卸载干净Oracle11g sophia天雪 orale数据库卸载干净清理注册表
完全卸载干净Oracle11g A、存在OUI卸载工具的情况下：第一步：停用所有Oracle相关的已启动的服务；第二步：找到OUI卸载工具：在“开始”菜单中找到“oracle_OraDb11g_home”文件夹中 &
apache 的access.log 日志文件太大如何解决 darkranger apache
CustomLog logs/access.log common 此写法导致日志数据一致自增变大。直接注释上面的语法 #CustomLog logs/access.log common 增加： CustomLog "|bin/rotatelogs.exe -l logs/access-%Y-%m-d.log
Hadoop单机模式环境搭建关键步骤 aijuans 分布式
Hadoop环境需要sshd服务一直开启，故，在服务器上需要按照ssh服务，以Ubuntu Linux为例，按照ssh服务如下： sudo apt-get install ssh sudo apt-get install rsync 编辑HADOOP_HOME/conf/hadoop-env.sh文件，将JAVA_HOME设置为Java
PL/SQL DEVELOPER 使用的一些技巧 atongyeye java sql
1 记住密码这是个有争议的功能，因为记住密码会给带来数据安全的问题。但假如是开发用的库，密码甚至可以和用户名相同，每次输入密码实在没什么意义，可以考虑让PLSQL Developer记住密码。位置：Tools菜单－－Preferences－－Oracle－－Logon HIstory－－Store with password 2 特殊Copy 在SQL Window
PHP：在对象上动态添加一个新的方法 bardo 方法动态添加闭包
有关在一个对象上动态添加方法，如果你来自Ruby语言或您熟悉这门语言，你已经知道它是什么...... Ruby提供给你一种方式来获得一个instancied对象，并给这个对象添加一个额外的方法。好！不说Ruby了，让我们来谈谈PHP PHP未提供一个“标准的方式”做这样的事情，这也是没有核心的一部分... 但无论如何，它并没有说我们不能做这样
ThreadLocal与线程安全 bijian1013 java java多线程 threadLocal
首先来看一下线程安全问题产生的两个前提条件： 1.数据共享，多个线程访问同样的数据。 2.共享数据是可变的，多个线程对访问的共享数据作出了修改。实例：定义一个共享数据： public static int a = 0;
Tomcat 架包冲突解决征客丶 tomcat Web
环境： Tomcat 7.0.6 win7 x64 错误表象：【我的冲突的架包是：catalina.jar 与 tomcat-catalina-7.0.61.jar 冲突，不知道其他架包冲突时是不是也报这个错误】严重: End event threw exception java.lang.NoSuchMethodException: org.apache.catalina.dep
【Scala三】分析Spark源代码总结的Scala语法一 bit1129 scala
Scala语法 1. classOf运算符 Scala中的classOf[T]是一个class对象，等价于Java的T.class,比如classOf[TextInputFormat]等价于TextInputFormat.class 2. 方法默认值 defaultMinPartitions就是一个默认值，类似C++的方法默认值
java 线程池管理机制 BlueSkator java线程池管理机制
编辑 Add Tools jdk线程池一、引言第一：降低资源消耗。通过重复利用已创建的线程降低线程创建和销毁造成的消耗。第二：提高响应速度。当任务到达时，任务可以不需要等到线程创建就能立即执行。第三：提高线程的可管理性。线程是稀缺资源，如果无限制的创建，不仅会消耗系统资源，还会降低系统的稳定性，使用线程池可以进行统一的分配，调优和监控。
关于hql中使用本地sql函数的问题（问-答） BreakingBad HQL 存储函数
转自于：http://www.iteye.com/problems/23775 问：我在开发过程中，使用hql进行查询（mysql5）使用到了mysql自带的函数find_in_set()这个函数作为匹配字符串的来讲效率非常好，但是我直接把它写在hql语句里面（from ForumMemberInfo fm,ForumArea fa where find_in_set(fm.userId,f
读《研磨设计模式》-代码笔记-迭代器模式-Iterator bylijinnan java 设计模式
声明：本文只为方便我个人查阅和理解，详细的分析以及源代码请移步原作者的博客http://chjavach.iteye.com/ import java.util.Arrays; import java.util.List; /** * Iterator模式提供一种方法顺序访问一个聚合对象中各个元素，而又不暴露该对象内部表示 * * 个人觉得，为了不暴露该
常用SQL chenjunt3 oracle sql C++c C#
--NC建库 CREATE TABLESPACE NNC_DATA01 DATAFILE 'E:\oracle\product\10.2.0\oradata\orcl\nnc_data01.dbf' SIZE 500M AUTOEXTEND ON NEXT 50M EXTENT MANAGEMENT LOCAL UNIFORM SIZE 256K ; CREATE TABLESPA
数学是科学技术的语言 comsci 工作活动领域模型
从小学到大学都在学习数学，从小学开始了解数字的概念和背诵九九表到大学学习复变函数和离散数学，看起来好像掌握了这些数学知识，但是在工作中却很少真正用到这些知识，为什么？最近在研究一种开源软件-CARROT2的源代码的时候，又一次感觉到数学在计算机技术中的不可动摇的基础作用，CARROT2是一种用于自动语言分类（聚类）的工具性软件，用JAVA语言编写，它
Linux系统手动安装rzsz 软件包 daizj linux sz rz
1、下载软件 rzsz-3.34.tar.gz。登录linux，用命令 wget http://freeware.sgi.com/source/rzsz/rzsz-3.48.tar.gz下载。 2、解压 tar zxvf rzsz-3.34.tar.gz 3、安装 cd rzsz-3.34 ; make posix 。注意：这个软件安装与常规的GNU软件不
读源码之:ArrayBlockingQueue dieslrae java
ArrayBlockingQueue是concurrent包提供的一个线程安全的队列,由一个数组来保存队列元素.通过 takeIndex和 putIndex来分别记录出队列和入队列的下标,以保证在出队列时不进行元素移动. //在出队列或者入队列的时候对takeIndex或者putIndex进行累加,如果已经到了数组末尾就又从0开始,保证数
C语言学习九枚举的定义和应用 dcj3sjt126com c
枚举的定义 # include <stdio.h> enum WeekDay { MonDay, TuesDay, WednesDay, ThursDay, FriDay, SaturDay, SunDay }; int main(void) { //int day; //day定义成int类型不合适 enum WeekDay day = Wedne
Vagrant 三种网络配置详解 dcj3sjt126com vagrant
Forwarded port Private network Public network Vagrant 中一共有三种网络配置，下面我们将会详解三种网络配置各自优缺点。端口映射(Forwarded port)，顾名思义是指把宿主计算机的端口映射到虚拟机的某一个端口上，访问宿主计算机端口时，请求实际是被转发到虚拟机上指定端口的。Vagrantfile中设定语法为： c
16.性能优化-完结 frank1234 性能优化
性能调优是一个宏大的工程，需要从宏观架构(比如拆分，冗余，读写分离，集群，缓存等)，软件设计（比如多线程并行化，选择合适的数据结构），数据库设计层面（合理的表设计，汇总表，索引，分区，拆分，冗余等）以及微观（软件的配置，SQL语句的编写，操作系统配置等）根据软件的应用场景做综合的考虑和权衡，并经验实际测试验证才能达到最优。性能水很深，笔者经验尚浅，赶脚也就了解了点皮毛而已，我觉得
Word Search hcx2013 search
Given a 2D board and a word, find if the word exists in the grid. The word can be constructed from letters of sequentially adjacent cell, where "adjacent" cells are those horizontally or ve
Spring4新特性——Web开发的增强 jinnianshilongnian spring spring mvc spring4
Spring4新特性——泛型限定式依赖注入 Spring4新特性——核心容器的其他改进 Spring4新特性——Web开发的增强 Spring4新特性——集成Bean Validation 1.1(JSR-349)到SpringMVC Spring4新特性——Groovy Bean定义DSL Spring4新特性——更好的Java泛型操作API Spring4新
CentOS安装配置tengine并设置开机启动 liuxingguome centos
yum install gcc-c++ yum install pcre pcre-devel yum install zlib zlib-devel yum install openssl openssl-devel Ubuntu上可以这样安装 sudo aptitude install libdmalloc-dev libcurl4-opens
第14章工具函数（上） onestopweb 函数
index.html <!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd"> <html xmlns="http://www.w3.org/
Xelsius 2008 and SAP BW at a glance blueoxygen BO Xelsius
Xelsius提供了丰富多样的数据连接方式，其中为SAP BW专属提供的是BICS。那么Xelsius的各种连接的优缺点比较以及Xelsius是如何直接连接到BEx Query的呢？以下Wiki文章应该提供了全面的概览。 http://wiki.sdn.sap.com/wiki/display/BOBJ/Xcelsius+2008+and+SAP+NetWeaver+BW+Co
oracle表空间相关 tongsh6 oracle
在oracle数据库中，一个用户对应一个表空间，当表空间不足时，可以采用增加表空间的数据文件容量，也可以增加数据文件，方法有如下几种： 1.给表空间增加数据文件 ALTER TABLESPACE "表空间的名字" ADD DATAFILE '表空间的数据文件路径' SIZE 50M; &nb
.Net framework4.0安装失败 yangjuanjava .net windows
上午的.net framework 4.0，各种失败，查了好多答案，各种不靠谱，最后终于找到答案了和Windows Update有关系，给目录名重命名一下再次安装，即安装成功了！下载地址：http://www.microsoft.com/en-us/download/details.aspx?id=17113 方法： 1.运行cmd，输入net stop WuAuServ 2.点击开