Small evaluation sets lie constantly
On forty cases, a jump from 70% to 80% is well within what chance produces. Teams ship prompt changes on differences like that regularly, and half of them are shipping noise.
The confidence interval is the number to read. If it spans zero, you do not have evidence of a difference, however encouraging the point estimate looks.
Significant is not the same as worth shipping
A one-point improvement can be real and still not worth a prompt that costs thirty percent more to run. Statistical significance answers "is it real"; it says nothing about "is it worth it".
Put both variants through the prompt diff to see the cost side before deciding.