Sentence Similarity
sentence-transformers
Safetensors
bert
feature-extraction
Generated from Trainer
dataset_size:278
loss:MatryoshkaLoss
loss:MultipleNegativesRankingLoss
Eval Results (legacy)
text-embeddings-inference
Instructions to use mahsaBa76/bge-base-custom-matryoshka with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use mahsaBa76/bge-base-custom-matryoshka with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("mahsaBa76/bge-base-custom-matryoshka") sentences = [ "How does Bitcoin's P2P network prevent malicious nodes from flooding the network with invalid blocks or transactions?", "paper-title: The Bitcoin Lightning Network: Scalable Off-Chain Instant Payments\n\n\\subsection*{8.4 Payment Routing}\nIt is theoretically possible to build a route map implicitly from observing 2 -of-2 multisigs on the blockchain to build a routing table. Note, however, this is not feasible with pay-to-script-hash transaction outputs, which can be resolved out-of-band from the bitcoin protocol via a third party routing service. Building a routing table will become necessary for large operators (e.g. BGP, Cjdns). Eventually, with optimizations, the network will look a lot like the correspondent banking network, or Tier-1 ISPs. Similar to how packets still reach their destination on your home network connection, not all participants need to have a full routing table. The core Tier-1 routes can be online all the time - while nodes at the edges, such as average users, would be connected intermittently.\n\nNode discovery can occur along the edges by pre-selecting and offering partial routes to well-known nodes.\n\n\\subsection*{8.5 Fees}\nLightning Network fees, which differ from blockchain fees, are paid directly between participants within the channel. The fees pay for the time-value of money for consuming the channel for a determined maximum period of time, and for counterparty risk of non-communication.\n\nCounterparty risk for fees only exist with one's direct channel counterparty. If a node two hops away decides to disconnect and their transaction gets broadcast on the blockchain, one's direct counterparties should not broadcast on the blockchain, but continue to update via novation with a new Commitment Transaction. See the Decrementing Timelocks entry in the HTLC section for more information about counterparty risk.\n\nThe time-value of fees pays for consuming time (e.g. 3 days) and is conceptually equivalent to a gold lease rate without custodial risk; it is the time-value for using up the access to money for a very short duration. Since certain paths may become very profitable in one direction, it is possible for fees to be negative to encourage the channel to be available for those profitable paths.\n\n\\section*{9 Risks}\nThe primary risks relate to timelock expiration. Additionally, for core nodes and possibly some merchants to be able to route funds, the keys must be held online for lower latency. However, end-users and nodes are able to keep their private keys firewalled off in cold storage.\n\n\\subsection*{9.1 Improper Timelocks}\nParticipants must choose timelocks with sufficient amounts of time. If insufficient time is given, it is possible that timelocked transactions believed to be invalid will become valid, enabling coin theft by the counterparty. There is a trade-off between longer timelocks and the time-value of money. When writing wallet and Lightning Network application software, it is necessary to ensure that sufficient time is given and users are able to have their transactions enter into the blockchain when interacting with non-cooperative or malicious channel counterparties.\n\n\\subsection*{9.2 Forced Expiration Spam}\nForced expiration of many transactions may be the greatest systemic risk when using the Lightning Network. If a malicious participant creates many channels and forces them all to expire at once, these may overwhelm block data capacity, forcing expiration and broadcast to the blockchain. The result would be mass spam on the bitcoin network. The spam may delay transactions to the point where other locktimed transactions become valid.\n\nThis may be mitigated by permitting one transaction replacement on all pending transactions. Anti-spam can be used by permitting only one transaction replacement of a higher sequence number by the inverse of an even or odd number. For example, if an odd sequence number was broadcast, permit a replacement to a higher even number only once. Transactions would use the sequence number in an orderly way to replace other transactions. This mitigates the risk assuming honest miners. This attack is extremely high risk, as incorrect broadcast of Commitment Transactions entail a full penalty of all funds in the channel.\n\nAdditionally, one may attempt to steal HTLC transactions by forcing a timeout transaction to go through when it should not. This can be easily mitigated by having each transfer inside the channel be lower than the total transaction fees used. Since transactions are extremely cheap and do not hit the blockchain with cooperative channel counterparties, large transfers of value can be split into many small transfers. This attempt can only work if the blocks are completely full for a long time. While it is possible to mitigate it using a longer HTLC timeout duration, variable block sizes may become common, which may need mitigations.\n\nIf this type of transaction becomes the dominant form of transactions which are included on the blockchain, it may become necessary to increase the block size and run a variable blocksize structure and timestop flags as described in the section below. This can create sufficient penalties and disincentives to be highly unprofitable and unsuccessful for attackers, as attackers lose all their funds from broadcasting the wrong transaction, to the point where it will never occur.", "paper-title: OmniLedger: A Secure, Scale-Out, Decentralized Ledger via Sharding\n\nFig. 11: Bootstrap bandwidth consumption with state blocks.\\\\[0pt]\nto create the UTXO state. For this experiment, we reconstructed Bitcoin's blockchain [5], [41] and created a parallel OmniLedger blockchain with weekly state blocks.\n\nFigure 11 depicts the bandwidth overhead of a validator that did not follow the state for the first 100 days. As we can see, the state block approach is better if the validator is outdated for more than 19 days or 2736 Bitcoin blocks.\n\nThe benefit might not seem substantial for Bitcoin, but in OmniLedger, 2736 blocks are created in less than 8 hours, meaning that for one day-long epochs, the state block approach is significantly better. If a peak throughput is required and 16 MB blocks are deployed, we expect reduced bandwidth consumption close to two orders of magnitude.\n\n\\section*{IX. Related Work}\nThe growing interests in scaling blockchains have produced a number of prominent systems that we compare in Table IV. ByzCoin [32] is a first step to scalable BFT consensus, but cannot scale-out. Elastico is the first open scale-out DL, however, it suffers from performance and security challenges that we have already discussed in Section II. RSCoin [16] proposes sharding as a scalable approach for centrally banked cryptocurrencies. RSCoin relies on a trusted source of randomness for sharding and auditing, making its usage problematic in trustless settings. Furthermore, to validate transactions, each shard has to coordinate with the client and instead of running BFT, RSCoin uses a simple two-phase commit, assuming that safety is preserved if the majority of validators is honest. This\n\nTABLE IV: Comparison of Distributed Ledger Systems\n\n\\begin{center}\n\\begin{tabular}{ccccccc}\n\\hline\nSystem & Scale-Out & \\begin{tabular}{c}\nCross-Shard \\\\\nTransaction Atomicity \\\\\n\\end{tabular} & State Blocks & \\begin{tabular}{c}\nMeasured Scalability \\\\\n(\\# of Validators) \\\\\n\\end{tabular} & \\begin{tabular}{c}\nEstimated \\\\\nTime to Fail \\\\\n\\end{tabular} & \\begin{tabular}{c}\nMeasured \\\\\nLatency \\\\\n\\end{tabular} \\\\\n\\hline\nRSCoin [16] & In Permissioned & Partial & No & 30 & N/A & 1 sec \\\\\nElastico [34] & In PoW & No & No & 1600 & 1 hour & 800 sec \\\\\nByzCoin [32] & No & N/A & No & 1008 & 19 years & 40 sec \\\\\nBitcoin-NG [21] & No & N/A & No & 1000 & N/A & 600 sec \\\\\nPBFT [9], [11] & No & N/A & No & 16 & N/A & 1 sec \\\\\nNakamoto [36] & No & N/A & No & 4000 & N/A & 600 sec \\\\\nOmniLedger & Yes & Yes & Yes & 2400 & 68.5 years & 1.5 sec \\\\\n\\hline\n\\end{tabular}\n\\end{center}\n\napproach, however, does not protect from double spending attempts by a malicious client colluding with a validator.\n\nIn short, prior solutions [16], [32], [34] achieve only two out of the three desired properties; decentralization, long-term security, and scale-out, as illustrated in Figure 1. OmniLedger overcomes this issue by scaling out, as far as throughput is concerned, and by maintaining consistency to the level required for safety, without imposing a total order.\n\nBitcoin-NG scales Bitcoin without changing the consensus algorithm by observing that the PoW process does not have to be the same as the transaction validation process; this results in two separate timelines: one slow for PoW and one fast for transaction validation. Although Bitcoin-NG significantly increases the throughput of Bitcoin, it is still susceptible to the same attacks as Bitcoin [24], [3].\n\nOther efforts to scale blockchains include: Tendermint [9], a protocol similar to PBFT for shard-level consensus that does not scale due to its similarities to PBFT, and the Lightning Network [40], an off-chain payment protocol for Bitcoin (also compatible to OmniLedger); it limits the amount of information committed to the blockchain.", "Datatype: lecture_note, Title: Lecture 4: Peer to Peer Networking for Blockchains\n\nHow does broadcast take only $O(\\log N)$ steps? We first need to understand the gossip-flooding-based broadcast protocol. The flooding protocol mimics the spread of an epidemic. Once a node is ``infected\", it infects its peers and forever stay's infected. It is easy to see that the spread of information will happen exponentially; hence the information will take $O(\\log N)$ hops to spread to all nodes. To formally understand the spread, we note that $d$-regular graphs with $d\\geq 3$ are an \\textit{expander graph} for large sizes ($|V|$) with high probability. An expander graph is a connected but sparse graph ($|E|=O(|V|)$) with the following property: $|\\partial A| \\geq \\epsilon|A|$ for any connected sub-graph $A$ with $|A|<0.5|V|$. Here, $|\\partial A|$ refers to the number of vertices outside $A$ with at least one neighbor in $A$. A gossip message originates with $A(0)$ as the broadcasting node with $|A(0)|=1$, in the next hop, it will spread to $\\partial A(0)$ with $|A(1)|\\geq (1+\\epsilon)|A(0)|$. This recursion continues and we have $|A(k)|\\geq(1+\\epsilon)^kA(0)$. Thus, the number of steps to reach half the number of nodes is logarithmic in the number of nodes. It can be shown that the other half of the nodes can also be covered in $O(\\log N)$ time.\n\n\n%Engineering issues (peer discovery, bootstrap, churn). Implementation connections (to the lab experiment). Validation of tx, blocks. How does that impact networking? What about skipping validation and doing cut-through routing? Compact blocks. (RR)\n\n\\section*{Bitcoin P2P network: A systems view}\nIn Bitcoin, peers connect to each other and communicate using the TCP protocol. The codebase allows for eight outgoing connections and up to 117 incoming connections. The network has a high churn rate (rate at which users enter/leave the system); hence, the node must be ready to connect to new peers. Moreover, to ensure that the peers we are connecting to are chosen randomly, the node keeps a large list of nodes running Bitcoin in the form of their (IP, port) tuple and establishes a connection to one of them randomly when a slot opens up. \n\nHow does a node bootstrap its list of peers? This happens by connecting to a set of DNS seed nodes. The seed nodes are not heavily decentralized; hence completely relying on the peer list provided by them is not advisable. On connecting to the initial set of peers, a node asks its neighbors for their peer list using {\\tt getAddr} and {\\tt Addr} messages. The node keeps refreshing its peer list regularly by exchanging peer lists with its peers. \n\nTransmission of all block and transactions happen through the inventory message {\\tt inv}, on receiving an {\\tt inv} message the node checks if it has the block or the transaction in its local storage. If not, it sends the {\\tt getData} message to fetch those blocks and transactions from the peer. Since block sizes are relatively large, block transmission can optionally happen in 2 stages. On receiving the {\\tt inv} message, the node may ask for headers first using {\\tt getHeaders} and ask for complete blocks only if a header chain is established. This header-first block transmission increases queries but can decrease the net bandwidth usage. It may also prevent nodes from accepting PoW invalid blocks since the node can check from the header whether PoW is valid. \n\nWe saw in the previous lecture that some nodes might be malicious. A question that may arise is: what stops malicious nodes from flooding the network with invalid blocks and transactions (i.e., with invalid PoW and/or signatures)? Such flooding will saturate the network and increase transmission delay to unacceptable levels. Such an attack is prevented by a simple design decision, forward message to peers only after validating the message; i.e., a node sends an {\\tt inv} block message to its peers only after validating the block. If the adversary creates an invalid block, the block will not be propagated beyond one honest node. Additionally, nodes maintain their peers' reputation using some predefined heuristics; if a peer misbehaves (say by sending a transaction with invalid signatures), its reputation is downgraded and after a certain lower threshold is disconnected." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
Welcome to the community
The community tab is the place to discuss and collaborate with the HF community!