zrpc: bound concurrency, and close idle connections before the node does
Review follow-ups to 73a93f1. Bound in-flight calls. rpcclient's single sendPostHandler goroutine imposed an accidental ceiling of one concurrent RPC; removing it without putting anything in its place left no bound at all. grpc-go supplies none either -- this server sets no MaxConcurrentStreams, so the default is math.MaxUint32 -- and dragonxd answers RPC with 8 worker threads behind a 4096-deep queue, shared on the pool node with getblocktemplate. Overload would therefore surface as mining latency rather than as an error we could back off on. MaxConnsPerHost blocks the caller at the limit instead of dialling more, which is the backpressure wanted; MaxIdleConnsPerHost alone would only cap reuse and let us exceed the limit while churning connections. Default 8, matching the node's DEFAULT_HTTP_THREADS, tunable with -rpc-max-concurrent. Even 8 removes all of the head-of-line blocking this work set out to fix. IdleConnTimeout 90s -> 20s. dragonxd closes idle connections at 30s (DEFAULT_HTTP_SERVER_TIMEOUT, applied via evhttp_set_timeout and not overridden in DRAGONX.conf). At 90s we were always the second to close, so a request could be written into a connection the server had already sent a FIN for, and Go will not retry a POST once bytes are on the wire. Closing first removes the race. Reject a negative -rpc-timeout, which silently meant "unbounded", the same as the documented 0. The check has to run after flag.Parse(); it was initially placed before it and never fired. Also correct the coinsupply note: hush_coinsupply walks the block index back to genesis, loading each block from disk and memoising newcoins and zfunds into the CBlockIndex, so the first call pays for the whole chain and later ones are nearly free. It is not a UTXO-set scan, as the earlier comment claimed. The measured 48s/3s figures are unchanged. Verified: five concurrent GetLightdInfo calls all return grpc-status 0 with no errors, 46 blocks ingested, and the daemon holds 2 sockets to the node rather than one per request; a negative timeout exits 1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FU87LdsJZiZkfq1eXubpeo
This commit is contained in:
@@ -71,10 +71,25 @@ type Client struct {
|
||||
nextID int64
|
||||
}
|
||||
|
||||
// New returns a client for a dragonxd JSON-RPC endpoint. A timeout of 0 means
|
||||
// no timeout, which reproduces the old unbounded behaviour and should not be
|
||||
// used in production.
|
||||
func New(addr, user, pass string, timeout time.Duration) *Client {
|
||||
// New returns a client for a dragonxd JSON-RPC endpoint.
|
||||
//
|
||||
// timeout of 0 means no timeout, which reproduces the old unbounded behaviour
|
||||
// and should not be used in production.
|
||||
//
|
||||
// maxConcurrent bounds how many requests may be in flight at once. This is not
|
||||
// optional book-keeping: rpcclient's single goroutine imposed an accidental
|
||||
// ceiling of ONE, and removing it without putting anything in its place would
|
||||
// let a burst of gRPC handlers fan out arbitrarily wide. grpc-go places no
|
||||
// limit of its own here -- this server sets no MaxConcurrentStreams, so the
|
||||
// default is math.MaxUint32. dragonxd serves RPC with 8 worker threads
|
||||
// (DEFAULT_HTTP_THREADS) behind a 4096-deep work queue, and on the pool node
|
||||
// those threads are shared with getblocktemplate, so overload shows up as
|
||||
// queueing latency for mining rather than as an error we could back off on.
|
||||
// A small number still removes all of the head-of-line blocking.
|
||||
func New(addr, user, pass string, timeout time.Duration, maxConcurrent int) *Client {
|
||||
if maxConcurrent < 1 {
|
||||
maxConcurrent = 1
|
||||
}
|
||||
return &Client{
|
||||
url: "http://" + addr,
|
||||
user: user,
|
||||
@@ -87,9 +102,22 @@ func New(addr, user, pass string, timeout time.Duration) *Client {
|
||||
// dragonxd's HTTP server supports keep-alive. rpcclient set
|
||||
// Close=true and opened a fresh TCP connection per request,
|
||||
// which left hundreds of sockets in TIME_WAIT on a busy node.
|
||||
MaxIdleConns: 32,
|
||||
MaxIdleConnsPerHost: 32,
|
||||
IdleConnTimeout: 90 * time.Second,
|
||||
//
|
||||
// MaxConnsPerHost is the real concurrency bound: it BLOCKS a
|
||||
// caller once the limit is reached rather than dialling more,
|
||||
// which is the backpressure we want. MaxIdleConnsPerHost only
|
||||
// caps reuse, so on its own it would let us exceed the limit
|
||||
// and go back to churning connections.
|
||||
MaxConnsPerHost: maxConcurrent,
|
||||
MaxIdleConns: maxConcurrent,
|
||||
MaxIdleConnsPerHost: maxConcurrent,
|
||||
// Must stay BELOW dragonxd's own idle timeout, which is 30s
|
||||
// (DEFAULT_HTTP_SERVER_TIMEOUT in httpserver.h, applied via
|
||||
// evhttp_set_timeout and not overridden in DRAGONX.conf).
|
||||
// Whoever closes second loses a race against a FIN already in
|
||||
// flight, and Go will not retry a POST once bytes are on the
|
||||
// wire -- so we close first.
|
||||
IdleConnTimeout: 20 * time.Second,
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user