最近做订单系统,id 生成这事儿看着简单,真要整好,坑不少。尤其是高并发场景,一个破 uuid 都能把数据库索引干趴下。所以打算把雪花算法从原理到落地捋一遍,顺便聊聊 spring boot 里怎么封装、怎么处理时钟回拨、怎么压榨性能。这篇文章是我实际写代码的经验总结,不是那种教科书式的贴代码,想到哪儿写到哪儿,但技术点都是真刀真 枪的。
1. 业务场景与方案选择
需要全局唯一 id 的地方太多了,订单号、消息 id、链路追踪的 traceid,还有日志 id、支付流水号。这些场景要求高并发、全局唯一、最好趋势递增,而且生成过程不能成为瓶颈。
常见的生成方案:
- uuid:最大的问题是无序、太长。字符串乱序,数据库索引性能差,排序也难受。优点是本地生成,没网络开销。
- 数据库自增 id:数字有序,但依赖数据库,分布式环境下有单点问题,并发一高就是瓶颈。
- 号段模式:每次从数据库拿一批号,性能不错,但依赖数据库,重启可能丢号段。
- 雪花算法:不依赖外部组件,趋势递增,性能极高,缺点也明显——时钟回拨会产生重复 id,还需要维护 workerid。
综合来看,雪花算法是主流,但得针对时钟回拨做增强。下面从原理开始。
2. 雪花算法结构解析
雪花算法最早是 twitter 搞出来的,核心思想很简洁:把一个 64 位整数按位拆开,分别表示时间戳、机器 id 和序列号。
0 1 - 41 42 - 51 52 - 63 ┌─┬─────────────┬─────────────┬─────────────┐ │0│ 时间戳(ms) │ 机器id(10bit)│ 序列号(12bit)│ └─┴─────────────┴─────────────┴─────────────┘
- 符号位:永远是 0,保证 id 是正数。
- 时间戳:41 位毫秒时间戳。2^41 毫秒差不多 69 年,实际用的时候要设一个起始时间,比如从 2024-01-01 开始,这样可以用到 2093 年。
- 机器 id:10 位,可以拆成 5 位机房 + 5 位机器,支持 32 个机房,每个机房 32 台机器,总共 1024 个节点。
- 序列号:12 位,同一毫秒内最多生成 4096 个 id。序列号用完了就等下一毫秒。
这个设计的好处很明显:时间戳 + 机器 id + 序列号组合起来,只要时钟不回拨、机器 id 不冲突,全局唯一;整体趋势递增,因为毫秒递增,同一毫秒内序列号递增;纯内存运算,没网络开销。
3. spring boot 实现一个雪花生成器
先建一个基础版,把位运算核心逻辑写出来。
public class snowflakeidgenerator {
// 起始时间戳:2024-01-01 00:00:00
private static final long start_timestamp = 1704067200000l;
// 各部分位数
private static final long sequence_bits = 12l;
private static final long worker_id_bits = 10l;
// 最大值
private static final long sequence_mask = ~(-1l << sequence_bits);
private static final long worker_id_mask = ~(-1l << worker_id_bits);
// 位移量
private static final long worker_id_shift = sequence_bits;
private static final long timestamp_shift = sequence_bits + worker_id_bits;
private final long workerid;
private long lasttimestamp = -1l;
private long sequence = 0l;
public snowflakeidgenerator(long workerid) {
if (workerid < 0 || workerid > worker_id_mask) {
throw new illegalargumentexception("workerid must be between 0 and " + worker_id_mask);
}
this.workerid = workerid;
}
public synchronized long nextid() {
long currenttimestamp = system.currenttimemillis();
if (currenttimestamp < lasttimestamp) {
// 先留空,后面处理时钟回拨
throw new illegalstateexception("clock moved backwards");
}
if (currenttimestamp == lasttimestamp) {
sequence = (sequence + 1) & sequence_mask;
if (sequence == 0) {
currenttimestamp = waitnextmillis(currenttimestamp);
}
} else {
sequence = 0l;
}
lasttimestamp = currenttimestamp;
return ((currenttimestamp - start_timestamp) << timestamp_shift)
| (workerid << worker_id_shift)
| sequence;
}
private long waitnextmillis(long currenttimestamp) {
while (currenttimestamp <= lasttimestamp) {
currenttimestamp = system.currenttimemillis();
}
return currenttimestamp;
}
}
配置里加数据中心和机器 id。我习惯把 10 位拆成 5 位机房 + 5 位机器,最终 workerid = (datacenterid << 5) | machineid。
@component
@configurationproperties(prefix = "snowflake")
public class snowflakeproperties {
private long datacenterid = 1;
private long machineid = 1;
// getter/setter 省略
public long getworkerid() {
return (datacenterid << 5) | machineid;
}
}
配置类里生成 bean:
@configuration
public class idgeneratorconfig {
@bean
public snowflakeidgenerator snowflakeidgenerator(snowflakeproperties properties) {
return new snowflakeidgenerator(properties.getworkerid());
}
}
这样就能直接 @autowired 使用了。但还没处理时钟回拨,接下来是关键。
4. 时钟回拨问题处理
雪花算法最怕系统时间往回跳,一回到过去,同一毫秒内就可能生成重复 id。业界常见的招数有三招:等待、抛错、历史时间补偿。
4.1 等待
如果回拨时间不长(比如 5 秒内),就让线程睡一会儿,等系统时间追上来。
if (currenttimestamp < lasttimestamp) {
long offset = lasttimestamp - currenttimestamp;
if (offset <= 5000) {
try {
thread.sleep(offset * 2);
} catch (interruptedexception e) {
thread.currentthread().interrupt();
}
currenttimestamp = system.currenttimemillis();
while (currenttimestamp < lasttimestamp) {
currenttimestamp = system.currenttimemillis();
}
} else {
throw new illegalstateexception("clock moved backwards too much");
}
}
思路很直白,但回拨频繁的话,线程会被反复阻塞,请求堆积。
4.2 直接抛错
回拨超过阈值,直接抛异常,让上层决定是降级还是拒绝。
if (currenttimestamp < lasttimestamp) {
long offset = lasttimestamp - currenttimestamp;
if (offset > 100) {
throw new illegalstateexception("clock moved backwards. refusing to generate id for " + offset + " ms");
}
// 小回拨,继续处理
}
这种策略容易导致服务不可用,但能保证不生成重复 id。
4.3 历史时间补偿
这个思路比较巧妙。既然当前时钟回拨了,那就假装时间还在上次生成 id 的时候,序列号继续往上加,直到加满 4096 个为止。这样避免等待,也不抛错。
public synchronized long nextid() {
long currenttimestamp = system.currenttimemillis();
if (currenttimestamp < lasttimestamp) {
sequence = (sequence + 1) & sequence_mask;
if (sequence == 0) {
throw new illegalstateexception("sequence exhausted while clock moved backwards");
}
// 继续用 lasttimestamp,序列号递增
return ((lasttimestamp - start_timestamp) << timestamp_shift)
| (workerid << worker_id_shift)
| sequence;
}
// 正常流程...
}
但注意,这个方案只能补偿短时间回拨。如果回拨太久,序列号用尽就只能抛错了。所以生产上我更喜欢“等待 + 补偿”的组合:
public synchronized long nextid() {
long currenttimestamp = system.currenttimemillis();
long offset = lasttimestamp - currenttimestamp;
if (offset > 0) {
if (offset > 5000) {
throw new illegalstateexception("clock moved backwards, offset=" + offset + "ms");
}
try {
thread.sleep(offset + 1);
} catch (interruptedexception e) {
thread.currentthread().interrupt();
}
currenttimestamp = system.currenttimemillis();
}
if (currenttimestamp == lasttimestamp) {
sequence = (sequence + 1) & sequence_mask;
if (sequence == 0) {
currenttimestamp = waitnextmillis(currenttimestamp);
}
} else {
sequence = 0l;
}
lasttimestamp = currenttimestamp;
return ((currenttimestamp - start_timestamp) << timestamp_shift)
| (workerid << worker_id_shift)
| sequence;
}
另外,如果你们有 redis 或者 zookeeper,可以把最近生成 id 的时间戳存一份,本机时间小于全局最大时间戳时,直接拿全局时间戳用。不过这会引入网络开销,我一般不用。
5. 预生成 id 池与双缓冲
雪花算法本身延迟已经很低了,但 system.currenttimemillis() 是有系统调用开销的,再加上 synchronized 锁,高并发下还是会被拖死。有一个土办法:先把 id 批量生成好放在本地队列里,用的时候直接从队列取,一步到位。
public class idpool {
private final snowflakeidgenerator generator;
private final blockingqueue<long> queue = new linkedblockingqueue<>(10000);
private final executorservice executor = executors.newsinglethreadexecutor();
public idpool(snowflakeidgenerator generator) {
this.generator = generator;
fill();
}
private void fill() {
for (int i = 0; i < 10000; i++) {
queue.offer(generator.nextid());
}
}
public long nextid() throws interruptedexception {
if (queue.size() < 5000) {
executor.submit(this::fill); // 异步补充,注意防重
}
return queue.take();
}
}
生产环境别这么裸,至少加个 atomicboolean 防止重复提交填充任务。也可以用双缓冲:两个队列,一个消费一个后台填充,消费到一半就交换。
public class doublebufferidpool {
private final snowflakeidgenerator generator;
private final int buffersize;
private volatile queue<long> currentbuffer;
private queue<long> nextbuffer;
private final executorservice executor = executors.newsinglethreadexecutor();
public doublebufferidpool(snowflakeidgenerator generator, int buffersize) {
this.generator = generator;
this.buffersize = buffersize;
this.currentbuffer = new arraydeque<>();
this.nextbuffer = new arraydeque<>();
fillbuffer(currentbuffer);
}
private void fillbuffer(queue<long> buffer) {
for (int i = 0; i < buffersize; i++) {
buffer.offer(generator.nextid());
}
}
public long nextid() {
long id = currentbuffer.poll();
if (id == null) {
synchronized (this) {
if (currentbuffer.isempty()) {
queue<long> tmp = currentbuffer;
currentbuffer = nextbuffer;
nextbuffer = tmp;
fillbuffer(nextbuffer);
}
id = currentbuffer.poll();
}
} else {
if (currentbuffer.size() <= buffersize / 2 && nextbuffer.size() < buffersize) {
executor.submit(() -> fillbuffer(nextbuffer));
}
}
return id;
}
}
预生成会浪费一些 id,但相比性能提升,这点浪费值得。另外,批量接口可以一次返回多个 id,进一步减少调用次数。
6. 集成 mybatis-plus
mybatis-plus 默认的 assign_id 就是雪花算法,但我们可以换成自己的生成器,做到统一管理。只需要实现 identifiergenerator 接口。
@component
public class mybatisplusidgenerator implements identifiergenerator {
@autowired
private idpool idpool; // 或者直接用 snowflakeidgenerator
@override
public number nextid(object entity) {
return idpool.nextid();
}
}
实体类主键用 @tableid(type = idtype.assign_id) 就行。如果想生成字符串订单号,可以单独写个服务:
@service
public class orderidservice {
@autowired
private snowflakeidgenerator generator;
public string generateorderid() {
return "order" + new simpledateformat("yyyymmdd").format(new date())
+ generator.nextid();
}
}
7. 美团 leaf 的设计思路
美团 leaf 是业界很经典的方案,两种模式都值得学习。
号段模式:从数据库拿一个号段(比如 1~1000),进程内按顺序分配。数据库表里有 biz_tag、max_id、step 这些字段。当号段快用完时,后台异步去数据库加载下一个号段。优点是对数据库压力小,id 趋势递增;缺点是要维护号段表,数据库挂了也玩不转。
雪花模式:leaf 做雪花模式的时候解决了两个痛点。一个是 workerid 动态分配,通过 zookeeper 注册节点,启动时获取递增序号作为 workerid,关闭时释放。另一个是时钟回拨,leaf 有 checktimestamp 机制,检测到回拨就抛异常,然后告警人工处理。
参考 leaf,我们在生产环境可以考虑用 redis 动态分配 workerid,但大多数团队规模用配置文件就行,只要不搞混。
8. 压测与监控
我拿 8c16g 的机器压了一次(jmeter 100 线程,60 秒),基础雪花算法 qps 大概 3.5 万,平均响应 2.8ms,p99 8.1ms。加了 id 池双缓冲之后,qps 能到 12 万,平均响应降到 0.8ms,p99 也只有 2.3ms。差距很明显。
生产环境建议重点监控这几个指标:
- 生成耗时:平均耗时、tp99。
- 池水位:剩余 id 数量,低于阈值就告警。
- 时钟回拨次数:记下每次回拨的时间差和频率。
- qps:评估容量和扩缩容。
用 micrometer + prometheus + grafana 就能搞定,代码里埋点记录一下 timer 和 counter 就行。
9. 总结
雪花算法不是什么银弹,但掌握了原理和变体,大部分场景都能应对。我的建议是:
- 起始时间戳设得远一点,保证 id 用几十年。
- workerid 别重复,运维要统一登记。
- 开启 ntp 同步,加
-x参数避免时间跳变。 - 日志里记下时钟回拨详情,出了问题好排查。
- id 生成服务要做限流降级,别让上游突起压垮。
以上就是springboot实现高并发分布式id生成的完整方案的详细内容,更多关于springboot高并发分布式id生成的资料请关注代码网其它相关文章!
发表评论