Downloads · 30 days
109
18% of all-time downloads
BoghdadyJR/al-MiniLM-L6-v2
al-MiniLM-L6-v2 is a sentence similarity model from BoghdadyJR. Use it when you need a score for how close two texts are. It is set up for sentence-transformers.
This is a sentence-transformers model finetuned from sentence-transformers/all-mpnet-base-v2 on the code-search-net/codesearchnet dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can…
Downloads · 30 days
109
18% of all-time downloads
All-time downloads
596
Public
Parameters
109M
438 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
This is a sentence-transformers model finetuned from sentence-transformers/all-mpnet-base-v2 on the code-search-net/code_search_net dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 384, 'do_lower_case': False}) with Transformer model: MPNetModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("BoghdadyJR/al-MiniLM-L6-v2")
# Run inference
sentences = [
'Keypoint.copy',
'def copy(self, x=None, y=None):\n """\n Create a shallow copy of the Keypoint object.\n\n Parameters\n ----------\n x : None or number, optional\n Coordinate of the keypoint on the x axis.\n If ``None``, the instance\'s value will be copied.\n\n y : None or number, optional\n Coordinate of the keypoint on the y axis.\n If ``None``, the instance\'s value will be copied.\n\n Returns\n -------\n imgaug.Keypoint\n Shallow copy.\n\n """\n return self.deepcopy(x=x, y=y)',
'def build_words_dataset(words=None, vocabulary_size=50000, printable=True, unk_key=\'UNK\'):\n """Build the words dictionary and replace rare words with \'UNK\' token.\n The most common word has the smallest integer id.\n\n Parameters\n ----------\n words : list of str or byte\n The context in list format. You may need to do preprocessing on the words, such as lower case, remove marks etc.\n vocabulary_size : int\n The maximum vocabulary size, limiting the vocabulary size. Then the script replaces rare words with \'UNK\' token.\n printable : boolean\n Whether to print the read vocabulary size of the given words.\n unk_key : str\n Represent the unknown words.\n\n Returns\n --------\n data : list of int\n The context in a list of ID.\n count : list of tuple and list\n Pair words and IDs.\n - count[0] is a list : the number of rare words\n - count[1:] are tuples : the number of occurrence of each word\n - e.g. [[\'UNK\', 418391], (b\'the\', 1061396), (b\'of\', 593677), (b\'and\', 416629), (b\'one\', 411764)]\n dictionary : dictionary\n It is `word_to_id` that maps word to ID.\n reverse_dictionary : a dictionary\n It is `id_to_word` that maps ID to word.\n\n Examples\n --------\n >>> words = tl.files.load_matt_mahoney_text8_dataset()\n >>> vocabulary_size = 50000\n >>> data, count, dictionary, reverse_dictionary = tl.nlp.build_words_dataset(words, vocabulary_size)\n\n References\n -----------------\n - `tensorflow/examples/tutorials/word2vec/word2vec_basic.py <https://github.com/tensorflow/tensorflow/blob/r0.7/tensorflow/examples/tutorials/word2vec/word2vec_basic.py>`__\n\n """\n if words is None:\n raise Exception("words : list of str or byte")\n\n count = [[unk_key, -1]]\n count.extend(collections.Counter(words).most_common(vocabulary_size - 1))\n dictionary = dict()\n for word, _ in count:\n dictionary[word] = len(dictionary)\n data = list()\n unk_count = 0\n for word in words:\n if word in dictionary:\n index = dictionary[word]\n else:\n index = 0 # dictionary[\'UNK\']\n unk_count += 1\n data.append(index)\n count[0][1] = unk_count\n reverse_dictionary = dict(zip(dictionary.values(), dictionary.keys()))\n if printable:\n tl.logging.info(\'Real vocabulary size %d\' % len(collections.Counter(words).keys()))\n tl.logging.info(\'Limited vocabulary size {}\'.format(vocabulary_size))\n if len(collections.Counter(words).keys()) < vocabulary_size:\n raise Exception(\n "len(collections.Counter(words).keys()) >= vocabulary_size , the limited vocabulary_size must be less than or equal to the read vocabulary_size"\n )\n return data, count, dictionary, reverse_dictionary',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
<!--
### Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details>
-->
<!--
### Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details>
-->
<!--
### Out-of-Scope Use
*List how the model may foreseeably be misused and address what users ought not to do with the model.*
-->
sts-dev| Metric | Value |
|---|---|
| pearson_cosine | 0.8806 |
| spearman_cosine | 0.881 |
| pearson_manhattan | 0.8781 |
| spearman_manhattan | 0.8798 |
| pearson_euclidean | 0.8794 |
| spearman_euclidean | 0.881 |
| pearson_dot | 0.8806 |
| spearman_dot | 0.881 |
| pearson_max | 0.8806 |
| spearman_max | 0.881 |
| func_name | whole_func_string | |
|---|---|---|
| type | string | string |
| details | <ul><li>min: 3 tokens</li><li>mean: 8.18 tokens</li><li>max: 21 tokens</li></ul> | <ul><li>min: 38 tokens</li><li>mean: 192.0 tokens</li><li>max: 384 tokens</li></ul> |
| func_name | whole_func_string |
|---|---|
| <code>ImageGraphCut.__msgc_step3_discontinuity_localization</code> | <code>def __msgc_step3_discontinuity_localization(self):<br> """<br> Estimate discontinuity in basis of low resolution image segmentation.<br> :return: discontinuity in low resolution<br> """<br> import scipy<br><br> start = self._start_time<br> seg = 1 - self.segmentation.astype(np.int8)<br> self.stats["low level object voxels"] = np.sum(seg)<br> self.stats["low level image voxels"] = np.prod(seg.shape)<br> # in seg is now stored low resolution segmentation<br> # back to normal parameters<br> # step 2: discontinuity localization<br> # self.segparams = sparams_hi<br> seg_border = scipy.ndimage.filters.laplace(seg, mode="constant")<br> logger.debug("seg_border: %s", scipy.stats.describe(seg_border, axis=None))<br> # logger.debug(str(np.max(seg_border)))<br> # logger.debug(str(np.min(seg_border)))<br> seg_border[seg_border != 0] = 1<br> logger.debug("seg_border: %s", scipy.stats.describe(seg_border, axis=None))<br> # scipy.ndimage.morphology.distance_transform_edt<br> boundary_dilatation_distance = self.segparams["boundary_dilatation_distance"]<br> seg = scipy.ndimage.morphology.binary_dilation(<br> seg_border,<br> # seg,<br> np.ones(<br> [<br> (boundary_dilatation_distance * 2) + 1,<br> (boundary_dilatation_distance * 2) + 1,<br> (boundary_dilatation_distance * 2) + 1,<br> ]<br> ),<br> )<br> if self.keep_temp_properties:<br> self.temp_msgc_lowres_discontinuity = seg<br> else:<br> self.temp_msgc_lowres_discontinuity = None<br><br> if self.debug_images:<br> import sed3<br><br> pd = sed3.sed3(seg_border) # ), contour=seg)<br> pd.show()<br> pd = sed3.sed3(seg) # ), contour=seg)<br> pd.show()<br> # segzoom = scipy.ndimage.interpolation.zoom(seg.astype('float'), zoom,<br> # order=0).astype('int8')<br> self.stats["t3"] = time.time() - start<br> return seg</code> |
| <code>ImageGraphCut.__multiscale_gc_lo2hi_run</code> | <code>def __multiscale_gc_lo2hi_run(self): # , pyed):<br> """<br> Run Graph-Cut segmentation with refinement of low resolution multiscale graph.<br> In first step is performed normal GC on low resolution data<br> Second step construct finer grid on edges of segmentation from first<br> step.<br> There is no option for use without use_boundary_penalties<br> """<br> # from PyQt4.QtCore import pyqtRemoveInputHook<br> # pyqtRemoveInputHook()<br> self._msgc_lo2hi_resize_init()<br> self.__msgc_step0_init()<br><br> hard_constraints = self.__msgc_step12_low_resolution_segmentation()<br> # ===== high resolution data processing<br> seg = self.__msgc_step3_discontinuity_localization()<br><br> self.stats["t3.1"] = (time.time() - self._start_time)<br> graph = Graph(<br> seg,<br> voxelsize=self.voxelsize,<br> nsplit=self.segparams["block_size"],<br> edge_weight_table=self._msgc_npenalty_table,<br> compute_low_nodes_index=True,<br> )<br><br> # graph.run() = graph.generate_base_grid() + graph.split_voxels()<br> # graph.run()<br> graph.generate_base_grid()<br> self.stats["t3.2"] = (time.time() - self._start_time)<br> graph.split_voxels()<br><br> self.stats["t3.3"] = (time.time() - self._start_time)<br><br> self.stats.update(graph.stats)<br> self.stats["t4"] = (time.time() - self._start_time)<br> mul_mask, mul_val = self.__msgc_tlinks_area_weight_from_low_segmentation(seg)<br> area_weight = 1<br> unariesalt = self.__create_tlinks(<br> self.img,<br> self.voxelsize,<br> self.seeds,<br> area_weight=area_weight,<br> hard_constraints=hard_constraints,<br> mul_mask=None,<br> mul_val=None,<br> )<br> # N-links prepared<br> self.stats["t5"] = (time.time() - self._start_time)<br> un, ind = np.unique(graph.msinds, return_index=True)<br> self.stats["t6"] = (time.time() - self._start_time)<br><br> self.stats["t7"] = (time.time() - self._start_time)<br> unariesalt2_lo2hi = np.hstack(<br> [unariesalt[ind, 0, 0].reshape(-1, 1), unariesalt[ind, 0, 1].reshape(-1, 1)]<br> )<br> nlinks_lo2hi = np.hstack([graph.edges, graph.edges_weights.reshape(-1, 1)])<br> if self.debug_images:<br> import sed3<br><br> ed = sed3.sed3(unariesalt[:, :, 0].reshape(self.img.shape))<br> ed.show()<br> import sed3<br><br> ed = sed3.sed3(unariesalt[:, :, 1].reshape(self.img.shape))<br> ed.show()<br> # ed = sed3.sed3(seg)<br> # ed.show()<br> # import sed3<br> # ed = sed3.sed3(graph.data)<br> # ed.show()<br> # import sed3<br> # ed = sed3.sed3(graph.msinds)<br> # ed.show()<br><br> # nlinks, unariesalt2, msinds = self.__msgc_step45678_construct_graph(area_weight, hard_constraints, seg)<br> # self.__msgc_step9_finish_perform_gc_and_reshape(nlinks, unariesalt2, msinds)<br> self.__msgc_step9_finish_perform_gc_and_reshape(<br> nlinks_lo2hi, unariesalt2_lo2hi, graph.msinds<br> )<br> self._msgc_lo2hi_resize_clean_finish()</code> |
| <code>ImageGraphCut.__multiscale_gc_hi2lo_run</code> | <code>def __multiscale_gc_hi2lo_run(self): # , pyed):<br> """<br> Run Graph-Cut segmentation with simplifiyng of high resolution multiscale graph.<br> In first step is performed normal GC on low resolution data<br> Second step construct finer grid on edges of segmentation from first<br> step.<br> There is no option for use without use_boundary_penalties<br> """<br> # from PyQt4.QtCore import pyqtRemoveInputHook<br> # pyqtRemoveInputHook()<br><br> self.__msgc_step0_init()<br> hard_constraints = self.__msgc_step12_low_resolution_segmentation()<br> # ===== high resolution data processing<br> seg = self.__msgc_step3_discontinuity_localization()<br> nlinks, unariesalt2, msinds = self.__msgc_step45678_hi2lo_construct_graph(<br> hard_constraints, seg<br> )<br> self.__msgc_step9_finish_perform_gc_and_reshape(nlinks, unariesalt2, msinds)</code> |
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}
| func_name | whole_func_string | |
|---|---|---|
| type | string | string |
| details | <ul><li>min: 3 tokens</li><li>mean: 9.23 tokens</li><li>max: 24 tokens</li></ul> | <ul><li>min: 50 tokens</li><li>mean: 276.31 tokens</li><li>max: 384 tokens</li></ul> |
| func_name | whole_func_string |
|---|---|
| <code>learn</code> | <code>def learn(env,<br> network,<br> seed=None,<br> lr=5e-4,<br> total_timesteps=100000,<br> buffer_size=50000,<br> exploration_fraction=0.1,<br> exploration_final_eps=0.02,<br> train_freq=1,<br> batch_size=32,<br> print_freq=100,<br> checkpoint_freq=10000,<br> checkpoint_path=None,<br> learning_starts=1000,<br> gamma=1.0,<br> target_network_update_freq=500,<br> prioritized_replay=False,<br> prioritized_replay_alpha=0.6,<br> prioritized_replay_beta0=0.4,<br> prioritized_replay_beta_iters=None,<br> prioritized_replay_eps=1e-6,<br> param_noise=False,<br> callback=None,<br> load_path=None,<br> **network_kwargs<br> ):<br> """Train a deepq model.<br><br> Parameters<br> -------<br> env: gym.Env<br> environment to train on<br> network: string or a function<br> neural network to use as a q function approximator. If string, has to be one of the names of registered models in baselines.common.models<br> (mlp, cnn, conv_only). If a function, should take an observation tensor and return a latent variable tensor, which<br> will be mapped to the Q function heads (see build_q_func in baselines.deepq.models for details on that)<br> seed: int or None<br> prng seed. The runs with the same seed "should" give the same results. If None, no seeding is used.<br> lr: float<br> learning rate for adam optimizer<br> total_timesteps: int<br> number of env steps to optimizer for<br> buffer_size: int<br> size of the replay buffer<br> exploration_fraction: float<br> fraction of entire training period over which the exploration rate is annealed<br> exploration_final_eps: float<br> final value of random action probability<br> train_freq: int<br> update the model every train_freq steps.<br> set to None to disable printing<br> batch_size: int<br> size of a batched sampled from replay buffer for training<br> print_freq: int<br> how often to print out training progress<br> set to None to disable printing<br> checkpoint_freq: int<br> how often to save the model. This is so that the best version is restored<br> at the end of the training. If you do not wish to restore the best version at<br> the end of the training set this variable to None.<br> learning_starts: int<br> how many steps of the model to collect transitions for before learning starts<br> gamma: float<br> discount factor<br> target_network_update_freq: int<br> update the target network every target_network_update_freq steps.<br> prioritized_replay: True<br> if True prioritized replay buffer will be used.<br> prioritized_replay_alpha: float<br> alpha parameter for prioritized replay buffer<br> prioritized_replay_beta0: float<br> initial value of beta for prioritized replay buffer<br> prioritized_replay_beta_iters: int<br> number of iterations over which beta will be annealed from initial value<br> to 1.0. If set to None equals to total_timesteps.<br> prioritized_replay_eps: float<br> epsilon to add to the TD errors when updating priorities.<br> param_noise: bool<br> whether or not to use parameter space noise (https://arxiv.org/abs/1706.01905)<br> callback: (locals, globals) -> None<br> function called at every steps with state of the algorithm.<br> If callback returns true training stops.<br> load_path: str<br> path to load the model from. (default: None)<br> **network_kwargs<br> additional keyword arguments to pass to the network builder.<br><br> Returns<br> -------<br> act: ActWrapper<br> Wrapper over act function. Adds ability to save it and load it.<br> See header of baselines/deepq/categorical.py for details on the act function.<br> """<br> # Create all the functions necessary to train the model<br><br> sess = get_session()<br> set_global_seeds(seed)<br><br> q_func = build_q_func(network, **network_kwargs)<br><br> # capture the shape outside the closure so that the env object is not serialized<br> # by cloudpickle when serializing make_obs_ph<br><br> observation_space = env.observation_space<br> def make_obs_ph(name):<br> return ObservationInput(observation_space, name=name)<br><br> act, train, update_target, debug = deepq.build_train(<br> make_obs_ph=make_obs_ph,<br> q_func=q_func,<br> num_actions=env.action_space.n,<br> optimizer=tf.train.AdamOptimizer(learning_rate=lr),<br> gamma=gamma,<br> grad_norm_clipping=10,<br> param_noise=param_noise<br> )<br><br> act_params = {<br> 'make_obs_ph': make_obs_ph,<br> 'q_func': q_func,<br> 'num_actions': env.action_space.n,<br> }<br><br> act = ActWrapper(act, act_params)<br><br> # Create the replay buffer<br> if prioritized_replay:<br> replay_buffer = PrioritizedReplayBuffer(buffer_size, alpha=prioritized_replay_alpha)<br> if prioritized_replay_beta_iters is None:<br> prioritized_replay_beta_iters = total_timesteps<br> beta_schedule = LinearSchedule(prioritized_replay_beta_iters,<br> initial_p=prioritized_replay_beta0,<br> final_p=1.0)<br> else:<br> replay_buffer = ReplayBuffer(buffer_size)<br> beta_schedule = None<br> # Create the schedule for exploration starting from 1.<br> exploration = LinearSchedule(schedule_timesteps=int(exploration_fraction * total_timesteps),<br> initial_p=1.0,<br> final_p=exploration_final_eps)<br><br> # Initialize the parameters and copy them to the target network.<br> U.initialize()<br> update_target()<br><br> episode_rewards = [0.0]<br> saved_mean_reward = None<br> obs = env.reset()<br> reset = True<br><br> with tempfile.TemporaryDirectory() as td:<br> td = checkpoint_path or td<br><br> model_file = os.path.join(td, "model")<br> model_saved = False<br><br> if tf.train.latest_checkpoint(td) is not None:<br> load_variables(model_file)<br> logger.log('Loaded model from {}'.format(model_file))<br> model_saved = True<br> elif load_path is not None:<br> load_variables(load_path)<br> logger.log('Loaded model from {}'.format(load_path))<br><br><br> for t in range(total_timesteps):<br> if callback is not None:<br> if callback(locals(), globals()):<br> break<br> # Take action and update exploration to the newest value<br> kwargs = {}<br> if not param_noise:<br> update_eps = exploration.value(t)<br> update_param_noise_threshold = 0.<br> else:<br> update_eps = 0.<br> # Compute the threshold such that the KL divergence between perturbed and non-perturbed<br> # policy is comparable to eps-greedy exploration with eps = exploration.value(t).<br> # See Appendix C.1 in Parameter Space Noise for Exploration, Plappert et al., 2017<br> # for detailed explanation.<br> update_param_noise_threshold = -np.log(1. - exploration.value(t) + exploration.value(t) / float(env.action_space.n))<br> kwargs['reset'] = reset<br> kwargs['update_param_noise_threshold'] = update_param_noise_threshold<br> kwargs['update_param_noise_scale'] = True<br> action = act(np.array(obs)[None], update_eps=update_eps, **kwargs)[0]<br> env_action = action<br> reset = False<br> new_obs, rew, done, _ = env.step(env_action)<br> # Store transition in the replay buffer.<br> replay_buffer.add(obs, action, rew, new_obs, float(done))<br> obs = new_obs<br><br> episode_rewards[-1] += rew<br> if done:<br> obs = env.reset()<br> episode_rewards.append(0.0)<br> reset = True<br><br> if t > learning_starts and t % train_freq == 0:<br> # Minimize the error in Bellman's equation on a batch sampled from replay buffer.<br> if prioritized_replay:<br> experience = replay_buffer.sample(batch_size, beta=beta_schedule.value(t))<br> (obses_t, actions, rewards, obses_tp1, dones, weights, batch_idxes) = experience<br> else:<br> obses_t, actions, rewards, obses_tp1, dones = replay_buffer.sample(batch_size)<br> weights, batch_idxes = np.ones_like(rewards), None<br> td_errors = train(obses_t, actions, rewards, obses_tp1, dones, weights)<br> if prioritized_replay:<br> new_priorities = np.abs(td_errors) + prioritized_replay_eps<br> replay_buffer.update_priorities(batch_idxes, new_priorities)<br><br> if t > learning_starts and t % target_network_update_freq == 0:<br> # Update target network periodically.<br> update_target()<br><br> mean_100ep_reward = round(np.mean(episode_rewards[-101:-1]), 1)<br> num_episodes = len(episode_rewards)<br> if done and print_freq is not None and len(episode_rewards) % print_freq == 0:<br> logger.record_tabular("steps", t)<br> logger.record_tabular("episodes", num_episodes)<br> logger.record_tabular("mean 100 episode reward", mean_100ep_reward)<br> logger.record_tabular("% time spent exploring", int(100 * exploration.value(t)))<br> logger.dump_tabular()<br><br> if (checkpoint_freq is not None and t > learning_starts and<br> num_episodes > 100 and t % checkpoint_freq == 0):<br> if saved_mean_reward is None or mean_100ep_reward > saved_mean_reward:<br> if print_freq is not None:<br> logger.log("Saving model due to mean reward increase: {} -> {}".format(<br> saved_mean_reward, mean_100ep_reward))<br> save_variables(model_file)<br> model_saved = True<br> saved_mean_reward = mean_100ep_reward<br> if model_saved:<br> if print_freq is not None:<br> logger.log("Restored model with mean reward: {}".format(saved_mean_reward))<br> load_variables(model_file)<br><br> return act</code> |
| <code>ActWrapper.save_act</code> | <code>def save_act(self, path=None):<br> """Save model to a pickle located at path"""<br> if path is None:<br> path = os.path.join(logger.get_dir(), "model.pkl")<br><br> with tempfile.TemporaryDirectory() as td:<br> save_variables(os.path.join(td, "model"))<br> arc_name = os.path.join(td, "packed.zip")<br> with zipfile.ZipFile(arc_name, 'w') as zipf:<br> for root, dirs, files in os.walk(td):<br> for fname in files:<br> file_path = os.path.join(root, fname)<br> if file_path != arc_name:<br> zipf.write(file_path, os.path.relpath(file_path, td))<br> with open(arc_name, "rb") as f:<br> model_data = f.read()<br> with open(path, "wb") as f:<br> cloudpickle.dump((model_data, self._act_params), f)</code> |
| <code>nature_cnn</code> | <code>def nature_cnn(unscaled_images, **conv_kwargs):<br> """<br> CNN from Nature paper.<br> """<br> scaled_images = tf.cast(unscaled_images, tf.float32) / 255.<br> activ = tf.nn.relu<br> h = activ(conv(scaled_images, 'c1', nf=32, rf=8, stride=4, init_scale=np.sqrt(2),<br> **conv_kwargs))<br> h2 = activ(conv(h, 'c2', nf=64, rf=4, stride=2, init_scale=np.sqrt(2), **conv_kwargs))<br> h3 = activ(conv(h2, 'c3', nf=64, rf=3, stride=1, init_scale=np.sqrt(2), **conv_kwargs))<br> h3 = conv_to_fc(h3)<br> return activ(fc(h3, 'fc1', nh=512, init_scale=np.sqrt(2)))</code> |
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}
eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16learning_rate: 2e-05num_train_epochs: 1warmup_ratio: 0.1fp16: Truebatch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonelearning_rate: 2e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 1max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Truefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falsebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | Training Loss | loss | sts-dev_spearman_cosine |
|---|---|---|---|---|
| 0 | 0 | - | - | 0.8810 |
| 0.08 | 100 | 0.4124 | 0.2191 | - |
| 0.16 | 200 | 0.108 | 0.0993 | - |
| 0.24 | 300 | 0.127 | 0.0756 | - |
| 0.32 | 400 | 0.0728 | - | - |
| 0.08 | 100 | 0.0662 | 0.0683 | - |
| 0.16 | 200 | 0.0321 | 0.0660 | - |
| 0.24 | 300 | 0.0815 | 0.0584 | - |
| 0.32 | 400 | 0.049 | 0.0591 | - |
| 0.4 | 500 | 0.0636 | 0.0612 | - |
| 0.48 | 600 | 0.0929 | 0.0577 | - |
| 0.56 | 700 | 0.0342 | 0.0568 | - |
| 0.64 | 800 | 0.0265 | 0.0572 | - |
| 0.72 | 900 | 0.0406 | 0.0551 | - |
| 0.8 | 1000 | 0.039 | 0.0549 | - |
| 0.88 | 1100 | 0.0376 | 0.0551 | - |
| 0.96 | 1200 | 0.0823 | 0.0556 | - |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
<!--
## Glossary
*Clearly define terms in order to be accessible across audiences.*
-->
<!--
## Model Card Authors
*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
-->
<!--
## Model Card Contact
*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
-->