Maximum_Likelihood

Statistics with Prof. Liu Sept 6, 2024

Statistics has two streams: frequenties, and Bayes.

Maximum likelihood is frequenties method.

Likelihood function is powerful, which contains all information for the data. We don't need others, just this function.

Likelihood function is the joint probability of all data. Likelihood 就是data的概率。所有data的信息全在这个likelihood finction里面了!

We assume data are iid. In statistics we say iid is "random sample". It means, e.g. 每次抽球的概率一样。

That Joint probability is just the product of all P(data | parameter)

It's data's probability! Not parameters' probability.

例:两个盒子,一个有5个黑球5个白球,另一个有9个黑球1个白球。现抽4次,每次放回地抽1个球。4个都是黑球。问最可能从哪个盒子抽的。

For box1, if 4 times, P(4 black) = P(black)^4 = 0.5^4

For box2, if 4 times, P(4 black) = P(black)^4 = 0.9^4 that's why iid and we make product.

The latter probability is larger, so we choose box2.

The parameters for each box: box1, binomial, n=4, p=0.5. box2, n=4, p=0.9

例,无穷多个盒子,它们有黑球的比例是从0到1不等。抽4次,4个都是黑球。问从哪个盒子抽的。

Now the parameters, n=4, but don't know p. Want to know p, once know p, we then know which box.

**We still choose the box with the highest p. We choose the max P(data given p) ie P(data given box). **

But what is the probability that box2 is what I have done? What is the probability for box1?

But these numbers are not the probabilities for the two boxes! It's more intuitive to make decisions based on their probabilities, such as the prob of rain 40% and not rain 60%.

It has logics. We are choosing the parameter which can make the data to be most likely to stand.

Likelihood is just the probability, the probability of data. 0 to 1.

**应用到科学方法论,We can measure the distance of a theory to the real world data, ie, to examine a theory is good or bad, using likelihood. **


Statistics do inference: estimation and prediction.

Estimation has two categories: 1. Assume we know the population distribution, we just estimate its parameters. 2. We don't even know the population distribution.

After that, if we use our model to fit new data, then it's prediction.

**Prediction error is larger than estimation error. ** Estimatiin error is just RSS, the sum of square residuals. But for prediction, a new dataset will introduce new noise, and plus the model's original RSS.


In logistics regression, and linear regression, and linear discriminate analysis, the conditional class probabilities sum up to 1, and thus is posterior probability P(parameter given data). Just compare them directly. We can use Bayes optimal classifier. It's not related to likelihood.

相关推荐
Hotchip_MEMS3 小时前
当雾化器遇上MEMS:一场从交互到制造的全链路效率提升
人工智能·笔记·物联网·电脑·制造
是上好佳佳佳呀3 小时前
【深度学习|Day02】PyTorch 深度学习笔记(下):张量运算与自动微分
pytorch·笔记·深度学习
IT古董5 小时前
【MES学习笔记系列】05 - MES 数据库设计
笔记·学习·mes
不瘦80斤不改名6 小时前
05-vibe-coding-向agentic-engineering演进
人工智能·笔记·python·prompt
九硕智慧建筑一体化厂家6 小时前
直流照明|无尘风淋室照明,高均匀无频闪,适配洁净车间高频合规工况
大数据·人工智能·笔记·智慧城市
Tasiorh6 小时前
《SVM支持向量机》在数学建模中实际运用
算法·机器学习·支持向量机·数学建模·svm
LATASA6 小时前
【 从0到1构建 Agent Harness学习笔记】
网络·笔记·学习
FakeOccupational6 小时前
【电路笔记 仿真】SPICE仿真器软件 LTspice:绘制简单电路(新建图页+添加组件)并仿真+模型导入+SUBCKT (子电路)绘制与使用+特殊器件绘制
笔记
是上好佳佳佳呀8 小时前
【深度学习|Day01】PyTorch 深度学习笔记(上):框架认知与张量基础
pytorch·笔记·深度学习
x17388062738 小时前
目录遍历漏洞
笔记·目录遍历漏洞